Merge branch 'hybrid-cloud-dev' into hybrid-cloud/tutorial/red-hat-haystack-rag
@@ -274,3 +274,48 @@ From the root of the project:
|
||||
```bash
|
||||
sass --watch --style=compressed ./qdrant-landing/themes/qdrant/static/css/pages/marketing-landing.scss ./qdrant-landing/themes/qdrant/static/css/marketing-landing.css
|
||||
```
|
||||
|
||||
## SEO
|
||||
|
||||
### Structured data (Schema.org, JSON-LD)
|
||||
|
||||
Structured data is a standardized format for providing information about a page and classifying the page content. It is used by search engines to understand the content of the page and to display rich snippets in search results.
|
||||
|
||||
We use JSON-LD format for structured data. Data is stored in JSON files in the `/assets/schema` directory. If no specific schema is provided for a page, the default schema is used based on the page type as defined in the `qdrant-landing/themes/qdrant/layouts/partials/seo_schema.html` file.
|
||||
|
||||
To add specific schema to a specific page, use the `seo_schema` or `seo_schema_json` parameter in the front matter of content markdown files (directory `content`).
|
||||
|
||||
To add json directly to the page, use the `seo_schema` parameter. The value should be a JSON object.
|
||||
|
||||
Example:
|
||||
|
||||
```yaml
|
||||
seo_schema: {
|
||||
"@context": "https://schema.org",
|
||||
"@type": "Organization",
|
||||
"name": "Qdrant",
|
||||
"url": "https://qdrant.io",
|
||||
"logo": "https://qdrant.io/images/logo.png",
|
||||
"sameAs": [
|
||||
"https://www.linkedin.com/company/qdrant",
|
||||
"https://twitter.com/qdrant"
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
To add a path to a JSON files with schema data, use the `seo_schema_json` parameter. This parameter should contain a list of paths to JSON files.
|
||||
The path should be relative to the `qdrant-landing/assets` directory.
|
||||
|
||||
Example:
|
||||
|
||||
```yaml
|
||||
seo_schema_json:
|
||||
- schema/schema-organization.json
|
||||
- schema/product-schema.json
|
||||
```
|
||||
|
||||
If you want to add a new schema, create a new JSON file in the `qdrant-landing/assets/schema` directory and add the path to the `seo_schema_json` parameter.
|
||||
|
||||
When use `seo_schema` and `seo_schema_json` together, `seo_schema` will be used additionally to `seo_schema_json` adding the second <script> tag with the `seo_schema` value.
|
||||
|
||||
Use `seo_schema_json` if you want to reuse the same schema for multiple pages to avoid duplication and make it easier to maintain.
|
||||
@@ -0,0 +1,24 @@
|
||||
{
|
||||
"@type": "Article",
|
||||
"@id": "{{- .Permalink -}}#article",
|
||||
"name": "{{- .Params.title | htmlEscape -}}",
|
||||
"headline": "{{- .Params.title | htmlEscape -}}",
|
||||
"image": [
|
||||
"{{- if .Params.social_preview_image -}}{{- .Params.social_preview_image | absURL -}}{{- end -}}"
|
||||
],
|
||||
"url": "{{- .Permalink -}}",
|
||||
"description": {{ $description := printf "%s" (.Params.description | plainify | replaceRE "(\n)" "" | replaceRE "[^\\w\\s:\\[\\]{}\"]" "" | htmlEscape ) -}}"{{- $description -}}",
|
||||
"abstract": {{- $abstract := .Summary -}}{{- if .Params.description -}}{{- $abstract = .Params.description -}}{{- end -}}{{ $processedAbstract := printf "%s" ($abstract | plainify | replaceRE "(\n)" "" | replaceRE "[^\\w\\s:\\[\\]{}\"]" "" | htmlEscape ) -}}"{{- $processedAbstract -}}",
|
||||
"wordCount": "{{- .WordCount -}}",
|
||||
"datePublished": "{{ .Date }}" ,
|
||||
"dateModified": "{{ .Date }}",{{- if .Params.author -}}
|
||||
"author": {
|
||||
"@type": "Person",
|
||||
"name": "{{- .Params.author | htmlEscape -}}"
|
||||
}{{- else -}}
|
||||
"author": {
|
||||
"@type": "Organization",
|
||||
"name": "Qdrant",
|
||||
"url": "https://qdrant.tech/"
|
||||
}{{- end -}}
|
||||
}
|
||||
@@ -0,0 +1,43 @@
|
||||
{
|
||||
"@type": "Organization",
|
||||
"@id": "{{- .Permalink -}}#organization",
|
||||
"name": "Qdrant",
|
||||
"legalName" : "Qdrant Solutions GmbH",
|
||||
"url": "https://qdrant.tech",
|
||||
"email": "info@qdrant.com",
|
||||
"logo": "https://qdrant.tech/images/logo_with_text.png",
|
||||
"description" : "{{- .Site.Params.description | htmlEscape -}}",
|
||||
"keywords" : [ {{- if .Params.Keywords -}}{{- range .Params.Keywords -}}"{{ . }}", {{- end -}}{{- else if .Site.Params.Keywords -}}{{- range .Site.Params.Keywords -}}"{{ . }}", {{- end -}}{{- end -}} "Qdrant" ],
|
||||
"foundingDate": "2021",
|
||||
"founders": [
|
||||
{
|
||||
"@type": "Person",
|
||||
"name": "{{ .Site.Params.Author }}"
|
||||
}, {
|
||||
"@type": "Person",
|
||||
"name": "Andre Zayarni"
|
||||
}
|
||||
],
|
||||
"location": "Berlin, Germany",
|
||||
"address": {
|
||||
"@type": "PostalAddress",
|
||||
"streetAddress": "Chausseestraße 86",
|
||||
"addressLocality": "Berlin",
|
||||
"addressRegion": "Berlin",
|
||||
"postalCode": "10115",
|
||||
"addressCountry": "DE"
|
||||
},
|
||||
"contactPoint": {
|
||||
"@type": "ContactPoint",
|
||||
"contactType": "customer support",
|
||||
"telephone": "+49 3040797694",
|
||||
"email": "info@qdrant.com"
|
||||
},
|
||||
"sameAs": [
|
||||
"{{ .Site.Params.github }}",
|
||||
"{{ .Site.Params.discord }}",
|
||||
"{{ .Site.Params.youtube }}",
|
||||
"https://www.linkedin.com/company/qdrant/",
|
||||
"https://twitter.com/qdrant_engine"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,51 @@
|
||||
{
|
||||
"@type": "Product",
|
||||
"@id": "{{- .Permalink -}}#product",
|
||||
"brand": {
|
||||
"@id": "https://qdrant.tech",
|
||||
"@type": "Organization",
|
||||
"name": "Qdrant Vector Database"
|
||||
},
|
||||
"description" : "{{- .Site.Params.description | htmlEscape -}}",
|
||||
"keywords" : [ {{- if .Params.Keywords -}}{{- range .Params.Keywords -}}"{{ . }}", {{- end -}}{{- else if .Site.Params.Keywords -}}{{- range .Site.Params.Keywords -}}"{{ . }}", {{- end -}}{{- end -}} "Qdrant" ],
|
||||
"name": "Qdrant",
|
||||
"image": "{{- if .Params.social_preview_image -}}{{- .Params.social_preview_image | absURL -}}{{- end -}}",
|
||||
"offers": {
|
||||
"@type": "AggregateOffer",
|
||||
"lowPrice": "0",
|
||||
"offerCount": "3",
|
||||
"priceCurrency": "USD",
|
||||
"offers": [
|
||||
{
|
||||
"@type": "Offer",
|
||||
"priceSpecification": {
|
||||
"@type": "PriceSpecification",
|
||||
"price": "free",
|
||||
"priceCurrency": "USD",
|
||||
"name": "Community",
|
||||
"url": "https://qdrant.tech/documentation/quick-start/"
|
||||
}
|
||||
},
|
||||
{
|
||||
"@type": "Offer",
|
||||
"priceSpecification": {
|
||||
"@type": "PriceSpecification",
|
||||
"price": "from 25$",
|
||||
"priceCurrency": "USD",
|
||||
"name": "Managed Cloud",
|
||||
"url": "https://qdrant.to/cloud"
|
||||
}
|
||||
},
|
||||
{
|
||||
"@type": "Offer",
|
||||
"priceSpecification": {
|
||||
"@type": "PriceSpecification",
|
||||
"price": "on request",
|
||||
"priceCurrency": "USD",
|
||||
"name": "Enterprise",
|
||||
"url": "https://qdrant.to/contact-us"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -54,7 +54,7 @@ As embeddings are vectors, one can apply a simple function to calculate the simi
|
||||
So with similarity learning, all we need to do is provide pairs of correct questions and answers.
|
||||
And then, the model will learn to distinguish proper answers by the similarity of embeddings.
|
||||
|
||||
>If you want to learn more about similarity learning and applications, check out this [article](https://blog.qdrant.tech/neural-search-tutorial-3f034ab13adc) which might be an asset.
|
||||
>If you want to learn more about similarity learning and applications, check out this [article](https://qdrant.tech/documentation/tutorials/neural-search/) which might be an asset.
|
||||
|
||||
## Let's build
|
||||
|
||||
|
||||
@@ -27,7 +27,7 @@ set up your collections.
|
||||
|
||||
Previously, you had to send multiple requests to the Qdrant API to perform multiple non-related tasks. However, this
|
||||
can cause significant network overhead and slow down the process, especially if you have a poor connection speed.
|
||||
Fortunately, the [new batch search feature](https://blog.qdrant.tech/batch-vector-search-with-qdrant-8c4d598179d5) allows
|
||||
Fortunately, the [new batch search feature](https://qdrant.tech/documentation/concepts/search/#batch-search-api) allows
|
||||
you to avoid this issue. With just one API call, Qdrant will handle multiple search requests in the most efficient way
|
||||
possible. This means that you can perform multiple tasks simultaneously without having to worry about network overhead
|
||||
or slow performance.
|
||||
@@ -37,14 +37,13 @@ or slow performance.
|
||||
To make our application accessible to ARM users, we have compiled it specifically for that platform. If it is not
|
||||
compiled for ARM, the device will have to emulate it, which can slow down performance. To ensure the best possible
|
||||
experience for ARM users, we have created Docker images specifically for that platform. Keep in mind that using
|
||||
a limited set of processor instructions may affect the performance of your vector search. Therefore, [we have tested
|
||||
both ARM and non-ARM architectures using similar setups to understand the potential impact on performance
|
||||
](https://blog.qdrant.tech/qdrant-supports-arm-architecture-363e92aa5026).
|
||||
a limited set of processor instructions may affect the performance of your vector search. Therefore, we have tested
|
||||
both ARM and non-ARM architectures using similar setups to understand the potential impact on performance.
|
||||
|
||||
## Full-text filtering
|
||||
|
||||
Qdrant is a vector database that allows you to quickly search for the nearest neighbors. However, you may need to apply
|
||||
additional filters on top of the semantic search. Up until version 0.10, Qdrant only supported keyword filters. With the
|
||||
release of Qdrant 0.10, [you can now use full-text filters](https://blog.qdrant.tech/qdrant-introduces-full-text-filters-and-indexes-9a032fcb5fa)
|
||||
release of Qdrant 0.10, [you can now use full-text filters](https://qdrant.tech/documentation/concepts/filtering/#full-text-match)
|
||||
as well. This new filter type can be used on its own or in combination with other filter types to provide even more
|
||||
flexibility in your searches.
|
||||
|
||||
@@ -175,6 +175,6 @@ After building your RAG chatbot, you'll be able to evaluate its performance agai
|
||||
|
||||
## What’s next?
|
||||
|
||||
Have a RAG project you want to bring to life? Join our [Discord community](discord.gg/qdrant) where we’re always sharing tips and answering questions on vector search and retrieval.
|
||||
Have a RAG project you want to bring to life? Join our [Discord community](https://discord.gg/qdrant) where we’re always sharing tips and answering questions on vector search and retrieval.
|
||||
|
||||
Learn more about how to properly evaluate your RAG responses: [Evaluating Retrieval Augmented Generation - a framework for assessment](https://superlinked.com/vectorhub/evaluating-retrieval-augmented-generation-a-framework-for-assessment).
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
draft: false
|
||||
title: "Qdrant is Now Available on Azure Marketplace!"
|
||||
short_description: Discover the power of Qdrant on Azure Marketplace!
|
||||
description: Discover the power of Qdrant on Azure Marketplace! Get started today and streamline your operations with ease.
|
||||
preview_image: /blog/azure-marketplace/azure-marketplace.png
|
||||
date: 2024-03-26T10:30:00Z
|
||||
author: David Myriel
|
||||
featured: false
|
||||
weight: 0
|
||||
tags:
|
||||
- Qdrant
|
||||
- Azure Marketplace
|
||||
- Enterprise
|
||||
- Vector Database
|
||||
---
|
||||
|
||||
We're thrilled to announce that Qdrant is now [officially available on Azure Marketplace](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db), bringing enterprise-level vector search directly to Azure's vast community of users. This integration marks a significant milestone in our journey to make Qdrant more accessible and convenient for businesses worldwide.
|
||||
|
||||
> *With the landscape of AI being complex for most customers, Qdrant's ease of use provides an easy approach for customers' implementation of RAG patterns for Generative AI solutions and additional choices in selecting AI components on Azure,* - Tara Walker, Principal Software Engineer at Microsoft.
|
||||
|
||||
## Why Azure Marketplace?
|
||||
|
||||
[Azure Marketplace](https://azuremarketplace.microsoft.com/en-us/) is renowned for its robust ecosystem, trusted by millions of users globally. By listing Qdrant on Azure Marketplace, we're not only expanding our reach but also ensuring seamless integration with Azure's suite of tools and services. This collaboration opens up new possibilities for our users, enabling them to leverage the power of Azure alongside the capabilities of Qdrant.
|
||||
|
||||
> *Enterprises like Bosch can now use the power of Microsoft Azure to host Qdrant, unleashing unparalleled performance and massive-scale vector search. "With Qdrant, we found the missing piece to develop our own provider independent multimodal generative AI platform at enterprise scale,* - Jeremy Teichmann (AI Squad Technical Lead & Generative AI Expert), Daly Singh (AI Squad Lead & Product Owner) - Bosch Digital.
|
||||
|
||||
## Key Benefits for Users:
|
||||
|
||||
- **Rapid Application Development:** Deploying a cluster on Microsoft Azure via the Qdrant Cloud console only takes a few seconds and can scale up as needed, giving developers maximal flexibility for their production deployments.
|
||||
|
||||
- **Billion Vector Scale:** Seamlessly grow and handle large-scale datasets with billions of vectors by leveraging Qdrant's features like vertical and horizontal scaling or binary quantization with Microsoft Azure's scalable infrastructure.
|
||||
|
||||
- **Unparalleled Performance:** Qdrant is built to handle scaling challenges, high throughput, low latency, and efficient indexing. Written in Rust makes Qdrant fast and reliable even under high load. See benchmarks.
|
||||
|
||||
- **Versatile Applications:** From recommendation systems to similarity search, Qdrant's integration with Microsoft Azure provides a versatile tool for a diverse set of AI applications.
|
||||
|
||||
## Getting Started:
|
||||
|
||||
Ready to experience the benefits of Qdrant on Azure Marketplace? Getting started is easy:
|
||||
|
||||
1. **Visit the Azure Marketplace**: Navigate to [Qdrant's Marketplace listing](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db).
|
||||
2. **Deploy Qdrant**: Follow the simple deployment instructions to set up your instance.
|
||||
3. **Start Using Qdrant**: Once deployed, start exploring the [features and capabilities of Qdrant](https://qdrant.tech/documentation/concepts/) on Azure.
|
||||
4. **Read Documentation**: Read Qdrant's [Documentation](https://qdrant.tech/documentation/) and build demo apps using [Tutorials](https://qdrant.tech/documentation/tutorials/).
|
||||
|
||||
## Join Us on this Exciting Journey:
|
||||
|
||||
We're incredibly excited about this collaboration with Azure Marketplace and the opportunities it brings for our users. As we continue to innovate and enhance Qdrant, we invite you to join us on this journey towards greater efficiency, scalability, and success.
|
||||
|
||||
Ready to elevate your business with Qdrant? **Click the banner and get started today!**
|
||||
|
||||
[](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db)
|
||||
|
||||
### About Qdrant:
|
||||
|
||||
Qdrant is the leading, high-performance, scalable, open-source vector database and search engine, essential for building the next generation of AI/ML applications. Qdrant is able to handle billions of vectors, supports the matching of semantically complex objects, and is implemented in Rust for performance, memory safety, and scale.
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
title: "Response to CVE-2024-2221: Arbitrary file upload vulnerability"
|
||||
draft: false
|
||||
slug: cve-2024-2221-response
|
||||
short_description: Qdrant keeps your systems secure
|
||||
description: Upgrade your deployments to at least v1.8.0. Cloud deployments not materially affected.
|
||||
preview_image: /blog/cve-2024-2221/cve-2024-2221-response-social-preview.png
|
||||
|
||||
# social_preview_image: /blog/Article-Image.png # Optional image used for link previews
|
||||
# title_preview_image: /blog/Article-Image.png # Optional image used for blog post title
|
||||
# small_preview_image: /blog/Article-Image.png # Optional image used for small preview in the list of blog posts
|
||||
date: 2024-04-05T13:00:00-07:00
|
||||
author: Mike Jang
|
||||
featured: false
|
||||
tags:
|
||||
- cve
|
||||
- security
|
||||
weight: 0 # Change this weight to change order of posts
|
||||
# For more guidance, see https://github.com/qdrant/landing_page?tab=readme-ov-file#blog
|
||||
---
|
||||
|
||||
### Summary
|
||||
|
||||
A security vulnerability has been discovered in Qdrant affecting all versions
|
||||
prior to v1.8, described in [CVE-2024-2221](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2024-2221).
|
||||
The vulnerability allows an attacker to upload arbitrary files to the
|
||||
filesystem, which can be used to gain remote code execution.
|
||||
|
||||
The vulnerability does not materially affect Qdrant cloud deployments, as that
|
||||
filesystem is read-only and authentication is enabled by default. At worst,
|
||||
the vulnerability could be used by an authenticated user to crash a cluster,
|
||||
which is already possible, such as by uploading more vectors than can fit in RAM.
|
||||
|
||||
Qdrant has addressed the vulnerability in v1.8.3 and above with code that
|
||||
restricts file uploads to a folder dedicated to that purpose.
|
||||
|
||||
### Action
|
||||
|
||||
Check the current version of your Qdrant deployment. Upgrade if your deployment
|
||||
is not at least v1.8.3.
|
||||
|
||||
To confirm the version of your Qdrant deployment in the cloud or on your local
|
||||
or cloud system, run an API GET call, as described in the [Qdrant Quickstart
|
||||
guide](https://qdrant.tech/documentation/cloud/quickstart-cloud/#step-2-test-cluster-access).
|
||||
If your Qdrant deployment is local, you do not need an API key.
|
||||
|
||||
Your next step depends on how you installed Qdrant. For details, read the
|
||||
[Qdrant Installation](https://qdrant.tech/documentation/guides/installation/)
|
||||
guide.
|
||||
|
||||
#### If you use the Qdrant container or binary
|
||||
|
||||
Upgrade your deployment. Run the commands in the applicable section of the
|
||||
[Qdrant Installation](https://qdrant.tech/documentation/guides/installation/)
|
||||
guide. The default commands automatically pull the latest version of Qdrant.
|
||||
|
||||
#### If you use the Qdrant helm chart
|
||||
|
||||
If you’ve set up Qdrant on kubernetes using a helm chart, follow the README in
|
||||
the [qdrant-helm](https://github.com/qdrant/qdrant-helm/tree/main?tab=readme-ov-file#upgrading) repository.
|
||||
Make sure applicable configuration files point to version v1.8.3 or above.
|
||||
|
||||
#### If you use the Qdrant cloud
|
||||
|
||||
No action is required. This vulnerability does not materially affect you. However, we suggest that you upgrade your cloud deployment to the latest version.
|
||||
@@ -0,0 +1,59 @@
|
||||
---
|
||||
draft: false
|
||||
title: "Introducing FastLLM: Qdrant’s Revolutionary LLM"
|
||||
short_description: The most powerful LLM known to human...or LLM.
|
||||
description: Lightweight and open-source. Custom made for RAG and completely integrated with Qdrant.
|
||||
preview_image: /blog/fastllm-announcement/fastllm.png
|
||||
date: 2024-04-01T00:00:00Z
|
||||
author: David Myriel
|
||||
featured: false
|
||||
weight: 0
|
||||
tags:
|
||||
- Qdrant
|
||||
- FastEmbed
|
||||
- LLM
|
||||
- Vector Database
|
||||
---
|
||||
|
||||
Today, we're happy to announce that **FastLLM (FLLM)**, our lightweight Language Model tailored specifically for Retrieval Augmented Generation (RAG) use cases, has officially entered Early Access!
|
||||
|
||||
Developed to seamlessly integrate with Qdrant, **FastLLM** represents a significant leap forward in AI-driven content generation. Up to this point, LLM’s could only handle up to a few million tokens.
|
||||
|
||||
**As of today, FLLM offers a context window of 1 billion tokens.**
|
||||
|
||||
However, what sets FastLLM apart is its optimized architecture, making it the ideal choice for RAG applications. With minimal effort, you can combine FastLLM and Qdrant to launch applications that process vast amounts of data. Leveraging the power of Qdrant's scalability features, FastLLM promises to revolutionize how enterprise AI applications generate and retrieve content at massive scale.
|
||||
|
||||
> *“First we introduced [FastEmbed](https://github.com/qdrant/fastembed). But then we thought - why stop there? Embedding is useful and all, but our users should do everything from within the Qdrant ecosystem. FastLLM is just the natural progression towards a large-scale consolidation of AI tools.” Andre Zayarni, President & CEO, Qdrant*
|
||||
>
|
||||
|
||||
## Going Big: Quality & Quantity
|
||||
|
||||
Very soon, an LLM will come out with a context window so wide, it will completely eliminate any value a measly vector database can add.
|
||||
|
||||
***We know this. That’s why we trained our own LLM to obliterate the competition. Also, in case vector databases go under, at least we'll have an LLM left!***
|
||||
|
||||
As soon as we entered Series A, we knew it was time to ramp up our training efforts. FLLM was trained on 300,000 NVIDIA H100s connected by 5Tbps Infiniband. It took weeks to fully train the model, but our unified efforts produced the most powerful LLM known to human…..or LLM.
|
||||
|
||||
We don’t see how any other company can compete with FastLLM. Most of our competitors will soon be burning through graphics cards trying to get to the next best thing. But it is too late. By this time next year, we will have left them in the dust.
|
||||
|
||||
> ***“Everyone has an LLM, so why shouldn’t we? Let’s face it - the more products and features you offer, the more they will sign up. Sure, this is a major pivot…but life is all about being bold.”*** *David Myriel, Director of Product Education, Qdrant*
|
||||
>
|
||||
|
||||
## Extreme Performance
|
||||
|
||||
Qdrant’s R&D is proud to stand behind the most dramatic benchmark results. Across a range of standard benchmarks, FLLM surpasses every single model in existence. In the [Needle In A Haystack](https://github.com/gkamradt/LLMTest_NeedleInAHaystack) (NIAH) test, FLLM found the embedded text with 100% accuracy, always within blocks containing 1 billion tokens. We actually believe FLLM can handle more than a trillion tokens, but it’s quite possible that it is hiding its true capabilities.
|
||||
|
||||
FastLLM has a fine-grained mixture-of-experts architecture and a whopping 1 trillion total parameters. As developers and researchers delve into the possibilities unlocked by this new model, they will uncover new applications, refine existing solutions, and perhaps even stumble upon unforeseen breakthroughs. As of now, we're not exactly sure what problem FLLM is solving, but hey, it's got a lot of parameters!
|
||||
|
||||
> *Our customers ask us “What can I do with an LLM this extreme?” I don’t know, but it can’t hurt to build another RAG chatbot.” Kacper Lukawski, Senior Developer Advocate, Qdrant*
|
||||
>
|
||||
|
||||
## Get Started!
|
||||
|
||||
Don't miss out on this opportunity to be at the forefront of AI innovation. Join FastLLM's Early Access program now and embark on a journey towards AI-powered excellence!
|
||||
|
||||
Stay tuned for more updates and exciting developments as we continue to push the boundaries of what's possible with AI-driven content generation.
|
||||
|
||||
Happy Generating! 🚀
|
||||
|
||||
[Sign Up for Early Access](https://qdrant.to/cloud)
|
||||
@@ -0,0 +1,185 @@
|
||||
---
|
||||
draft: true
|
||||
title: Gen AI and Vector Search - Iveta Lohovska | Vector Space Talks
|
||||
slug: gen-ai-and-vector-search
|
||||
short_description: Iveta emphasizes the importance of trustworthy AI,
|
||||
particularly when implementing it within high-stakes enterprises like
|
||||
governments and security agencies
|
||||
description: Iveta Lohovska discusses the importance of explainability and
|
||||
transparency, discussing high-stakes use cases in sectors like cybersecurity
|
||||
and climate data, and emphasizing the necessity for on-prem solutions and
|
||||
traceable vector databases to ensure data integrity and confidentiality.
|
||||
preview_image: /blog/from_cms/iveta-lohovska-bp-cropped.png
|
||||
date: 2024-04-04T21:28:00.000Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
- Vector Space Talks
|
||||
- Vector Search
|
||||
- Retrieval Augmented Generation
|
||||
- GenAI
|
||||
---
|
||||
> *"In the generative AI context of AI, all foundational models have been trained on some foundational data sets that are distributed in different ways. Some are very conversational, some are very technical, some are on, let's say very strict taxonomy like healthcare or chemical structures. We call them modalities, and they have different representations.”*\
|
||||
— Iveta Lohovska
|
||||
>
|
||||
|
||||
Iveta Lohovska serves as the Chief Technologist and Principal Data Scientist for AI and Supercomputing at Hewlett Packard Enterprise (HPE), where she champions the democratization of decision intelligence and the development of ethical AI solutions. An industry leader, her multifaceted expertise encompasses natural language processing, computer vision, and data mining. Committed to leveraging technology for societal benefit, Iveta is a distinguished technical advisor to the United Nations' AI for Good program and a Data Science lecturer at the Vienna University of Applied Sciences. Her career also includes impactful roles with the World Bank Group, focusing on open data initiatives and Sustainable Development Goals (SDGs), as well as collaborations with USAID and the Gates Foundation.
|
||||
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/7f1RDwp5l2Ps9N7gKubl8S?si=kCSX4HGCR12-5emokZbRfw), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/RsRAUO-fNaA).***
|
||||
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/RsRAUO-fNaA?si=s3k_-DP1U0rkPlEV" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
<iframe src="https://podcasters.spotify.com/pod/show/qdrant-vector-space-talk/embed/episodes/Gen-AI-and-Vector-Search---Iveta-Lohovska--Vector-Space-Talks-020-e2hnie2/a-ab48uha" height="102px" width="400px" frameborder="0" scrolling="no"></iframe>
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
In our continuous pursuit of knowledge and understanding, especially in the evolving landscape of AI and the vector space, we brought another great Vector Space Talk episode featuring Iveta Lohovska as she talks about generative AI and vector search.
|
||||
|
||||
Iveta brings valuable insights from her work with the World Bank and as Chief Technologist at HPE, explaining the ins and outs of ethical AI implementation.
|
||||
|
||||
Here are the episode highlights:
|
||||
- Exploring the critical role of trustworthiness and explainability in AI, especially within high confidentiality use cases like government and security agencies.
|
||||
- Discussing the importance of transparency in AI models and how it impacts the handling of data and understanding the foundational datasets for vector search.
|
||||
- Iveta shares her experiences implementing generative AI in high-stakes environments, including the energy sector and policy-making, emphasizing accuracy and source credibility.
|
||||
- Strategies for managing data privacy in high-stakes sectors, the superiority of on-premises solutions for control, and the implications of opting for cloud or hybrid infrastructure.
|
||||
- Iveta's take on the maturity levels of generative AI, the ongoing development of smaller, more focused models, and the evolving landscape of AI model licensing and open-source contributions.
|
||||
|
||||
> Fun Fact: The climate agent solution showcased by Iveta helps individuals benchmark their carbon footprint and assists policymakers in drafting policy recommendations based on scientifically accurate data.
|
||||
>
|
||||
|
||||
## Show notes:
|
||||
|
||||
00:00 AI's vulnerabilities and ethical implications in practice.\
|
||||
06:28 Trust reliable sources for accurate climate data.\
|
||||
09:14 Vector database offers control and explainability.\
|
||||
13:21 On-prem vital for security and control.\
|
||||
16:47 Gen AI chat models at basic maturity.\
|
||||
19:28 Mature technical community, but slow enterprise adoption.\
|
||||
23:34 Advocates for open source but highlights complexities.\
|
||||
25:38 Unreliable information, triangle of necessities, vector space.
|
||||
|
||||
## More Quotes from Iveta:
|
||||
|
||||
*"What we have to ensure here is that every citation and every answer and augmentation by the generative AI on top of that is linked to the exact source of paper or publication, where it's coming from, to ensure that we can trace it back to where the climate information is coming from.”*\
|
||||
— Iveta Lohovska
|
||||
|
||||
*"Explainability means if you receive a certain answer based on your prompt, you can trace it back to the exact source where the embedding has been stored or the source of where the information is coming from and things.”*\
|
||||
— Iveta Lohovska
|
||||
|
||||
*"Chat GPT for conversational purposes and individual help is something very cool but when this needs to be translated into actual business use cases scenario with all the constraint of the enterprise architecture, with the constraint of the use cases, the reality changes quite dramatically.”*\
|
||||
— Iveta Lohovska
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
Look at that. We are back for another vector space talks. I'm very excited to be doing this today with you all. I am joined by none other than Sabrina again. Where are you at, Sabrina? How's it going?
|
||||
|
||||
Sabrina Aquino:
|
||||
Hey there, Demetrios. Amazing. Another episode and I'm super excited for this one. How are you doing?
|
||||
|
||||
Demetrios:
|
||||
I'm great. And we're going to bring out our guest of honor today. We are going to be talking a lot about trustworthy AI because Iveta has a background working with the World bank and focusing on the open data with that. But currently she is chief technologist and principal data scientist at HPE. And we were talking before we hit record before we went live. And we've got some hot takes that are coming up. So I'm going to bring Iveta to the stage. Where are you? There you are, our guest of honor.
|
||||
|
||||
Demetrios:
|
||||
How you doing?
|
||||
|
||||
Iveta Lohovska:
|
||||
Good. I hope you can hear me well.
|
||||
|
||||
Demetrios:
|
||||
Loud and clear. Yes.
|
||||
|
||||
Iveta Lohovska:
|
||||
Happy to join here from Vienna and thank you for the invite.
|
||||
|
||||
Demetrios:
|
||||
Yes. So I'm very excited to talk with you today. I think it's probably worth getting the TLDR on your story and why you're so passionate about trustworthiness and explainability.
|
||||
|
||||
Iveta Lohovska:
|
||||
Well, I think especially in the genaid context where if there any vulnerabilities around the solution or the training data set or any underlying context, either in the enterprise or in a smaller scale, it's just the scale that AI engine AI can achieve if it has any vulnerabilities or any weaknesses when it comes to explainability or trustworthiness or bias, it just goes explain nature. So it is to be considered and taken with high attention when it comes to those use cases. And most of my work is within an enterprise with high confidentiality use cases. So it plays a big role more than actually people will think it's on a high level. It just sounds like AI ethical principles or high level words that are very difficult to implement in technical terms. But in reality, when you hit the ground, when you hit the projects, when you work with in the context of, let's say, governments or organizations that deal with atomic energy, I see it in Vienna, the atomic agency is a neighboring one, or security agencies. Then you see the importance and the impact of those terms and the technical implications behind that.
|
||||
|
||||
Sabrina Aquino:
|
||||
That's amazing. And can you talk a little bit more about the importance of the transparency of these models and what can happen if we don't know exactly what kind of data they are being trained on?
|
||||
|
||||
Iveta Lohovska:
|
||||
I mean, this is especially relevant under our context of vector databases and vector search. Because in the generative AI context of AI, all foundational models have been trained on some foundational data sets that are distributed in different ways. Some are very conversational, some are very technical, some are on, let's say very strict taxonomy like healthcare or chemical structures. We call them modalities, and they have different representations. So, so when it comes to implementing vector search or vector database and knowing the distribution of the foundational data sets, you have better control if you introduce additional layers or additional components to have the control in your hands of where the information is coming from, where it's stored, what are the embeddings. So that helps, but it is actually quite important that you know what the foundational data sets are, so that you can predict any kind of weaknesses or vulnerabilities or penetrations that the solution or the use case of the model will face when it lands at the end user. Because we know with generative AI that is unpredictable, we know we can implement guardrails. They're already solutions.
|
||||
|
||||
Iveta Lohovska:
|
||||
We know they're not 100, they don't give you 100% certainty, but they are definitely use cases and work where you need to hit the hundred percent certainty, especially intelligence, cybersecurity and healthcare.
|
||||
|
||||
Demetrios:
|
||||
Yeah, that's something that I wanted to dig into a little bit. More of these high stakes use cases feel like you can't. I don't know. I talk with a lot of people about at this current time, it's very risky to try and use specifically generative AI for those high stakes use cases. Have you seen people that are doing it well, and if so, how?
|
||||
|
||||
Iveta Lohovska:
|
||||
Yeah, I'm in the business of high stakes use cases and yes, we do those kind of projects and work, which is very exciting and interesting, and you can see the impact. So I'm in the generative AI implementation into enterprise control. An enterprise context could mean critical infrastructure, could mean telco, could mean a government, could mean intelligence organizations. So those are just a few examples, but I could flip the coin and give you an alternative for a public one where I can share, let's say a good example is climate data. And we recently worked on, on building a knowledge worker, a climate agent that is trained, of course, his foundational knowledge, because all foundational models have prior knowledge they can refer to. But the key point here is to be an expert on climate data emissions gap country cards. Every country has a commitment to meet certain reduction emission reduction goals and then benchmarked and followed through the international supervisions of the world, like the United nations environmental program and similar entities. So when you're training this agent on climate data, they're competing ideas or several sources.
|
||||
|
||||
Iveta Lohovska:
|
||||
You can source your information from the local government that is incentivized to show progress to the nation and other stakeholders faster than the actual reality, the independent entities that provide information around the state of the world when it comes to progress towards certain climate goals. And there are also different parties. So for this kind of solution, we were very lucky to work with kind of the status co provider, the benchmark around climate data, around climate publications. And what we have to ensure here is that every citation and every answer and augmentation by the generative AI on top of that is linked to the exact source of paper or publication, where it's coming from, to ensure that we can trace it back to where the climate information is coming from. If Germany performs better compared to Austria, and also the partner we work with was the United nations environmental program. So they want to make sure that they're the citadel scientific arm when it comes to giving information. And there's no compromise, could be a compromise on the structure of the answer, on the breadth and death of the information, but there should be no compromise on the exact fact fullness of the information and where it's coming from. And this is a concrete example because why, you oughta ask, why is this so important? Because it has two interfaces.
|
||||
|
||||
Iveta Lohovska:
|
||||
It has the public. You can go and benchmark your carbon footprint as an individual living in one country comparing to an individual living in another. But if you are a policymaker, which is the other interface of this application, who will write the policy recommendation of a country in their own country, or a country they're advising on, you might want to make sure that the scientific citations and the policy recommendations that you're making are correct and they are retrieved from the proper data sources. Because there will be a huge implication when you go public with those numbers or when you actually design a law that is reinforceable with legal terms and law enforcement.
|
||||
|
||||
Sabrina Aquino:
|
||||
That's very interesting, Iveta, and I think this is one of the great use cases for RAG, for example. And I think if you can talk a little bit more about how vector search is playing into all of this, how it's helping organizations do this, this.
|
||||
|
||||
Iveta Lohovska:
|
||||
Would be amazing in such specific use cases. I think the main differentiator is the traceability component, the first that you have full control on which data it will refer to, because if you deal with open source models, most of them are open, but the data it has been trained on has not been opened or given public so with vector database you introduce a step of control and explainability. Explainability means if you receive a certain answer based on your prompt, you can trace it back to the exact source where the embedding has been stored or the source of where the information is coming from and things. So this is a major use case for us for those kind of high stake solution is that you have the explainability and traceability. Explainability. It could be as simple as a semantical similarity to the text, but also the traceability of where it's coming from and the exact link of where it's coming from. So it should be, it shouldn't be referred. You can close and you can cut the line of the model referring to its previous knowledge by introducing a vector database, for example.
|
||||
|
||||
Iveta Lohovska:
|
||||
So there could be many other implications and improvements in terms of speed and just handling huge amounts of data, yet also nice to have that come with this kind of technique, but the prior use case is actually not incentivized around those.
|
||||
|
||||
Demetrios:
|
||||
So if I'm hearing you correctly, it's like yet another reason why you should be thinking about using vector databases, because you need that ability to cite your work and it's becoming a very strong design pattern. Right. We all understand now, if you can't see where this data has been pulled from or you can't get, you can't trace back to the actual source, it's hard to trust what the output is.
|
||||
|
||||
Iveta Lohovska:
|
||||
Yes, and the easiest way to kind of cluster the two groups. If you think of creative fields and marketing fields and design fields where you could go wild and crazy with the temperature on each model, how creative it could go and how much novelty it could bring to the answer are one family of use cases. But there is exactly the opposite type of use cases where this is a no go and you don't need any creativity, you just focus on, focus on the factfulness and explainability. So it's more of the speed and the accuracy of retrieving information with a high level of novelty, but not compromising on any kind of facts within the answer, because there will be legal implications and policy implications and societal implications based on the action taken on this answer, either policy recommendation or legal action. There's a lot to do with the intelligence agencies that retrieve information based on nearest neighbor or kind of a relational analysis that you can also execute with vector databases and generative AI.
|
||||
|
||||
Sabrina Aquino:
|
||||
And we know that for these high stakes sectors that data privacy is a huge concern. And when we're talking about using vector databases and storing that data somewhere, what are some of the principles or techniques that you use in terms of infrastructure, where should you store your vector database and how should you think about that part of your system?
|
||||
|
||||
Iveta Lohovska:
|
||||
Yeah, so most of the cases, I would say 99% of the cases, is that if you have such a high requirements around security and explainability, security of the data, but those security of the whole use case and environment, and the explainability and trustworthiness of the answer, then it's very natural to have expectations that will be on prem and not in the cloud, because only on prem you have a full control of where your data sits, where your model sits, the full ownership of your IP, and then the full ownership of having less question marks of the implementation and architecture, but mainly the full ownership of the end to end solution. So when it comes to those use cases, RAG on Prem, with the whole infrastructure, with the whole software and platform layers, including models on Prem, not accessible through an API, through a service somewhere where you don't know where the guardrails is, who designed the guardrails, what are the guardrails? And we see those, this a lot with, for example, copilot, a lot of question marks around that. So it's a huge part of my work is just talking of it, just sorting out that.
|
||||
|
||||
Sabrina Aquino:
|
||||
Exactly. You don't want to just give away your data to a cloud provider, because there's many implications that that comes with. And I think even your clients, they need certain certifications, then they need to make sure that nobody can access that data, something that you cannot. Exactly. I think ensure if you're just using a cloud provider somewhere, which is, I think something that's very important when you're thinking about these high stakes solutions. But also I think if you're going to maybe outsource some of the infrastructure, you also need to think about something that's similar to a hybrid cloud solution where you can keep your data and outsource the kind of management of infrastructure. So that's also a nice use case for that, right?
|
||||
|
||||
Iveta Lohovska:
|
||||
I mean, I work for HPE, so hybrid is like one of our biggest sacred words. Yeah, exactly. But actually like if you see the trends and if you see how expensive is to work to run some of those workloads in the cloud, either for training for national model or fine tuning. And no one talks about inference, inference not in ten users, but inference in hundred users with big organizations. This itself is not sustainable. Honestly, when you do the simple Linux, algebra or math of the exponential cost around this. That's why everything is hybrid. And there are use cases that make sense to be fast and speedy and easy to play with, low risk in the cloud to try.
|
||||
|
||||
Iveta Lohovska:
|
||||
But when it comes to actual GenAI work and LLM models, yeah, the answer is never straightforward when it comes to the infrastructure and the environment where you are hosting it, for many reasons, not just cost, but any other.
|
||||
|
||||
Demetrios:
|
||||
So there's something that I've been thinking about a lot lately that I would love to get your take on, especially because you deal with this day in and day out, and it is the maturity levels of the current state of Gen AI and where we are at for chat GPT or just llms and foundational models feel like they just came out. And so we're almost in the basic, basic, basic maturity levels. And when you work with customers, how do you like kind of signal that, hey, this is where we are right now, but you should be very conscientious that you're going to need to potentially work with a lot of breaking changes or you're going to have to be constantly updating. And this isn't going to be set it and forget it type of thing. This is going to be a lot of work to make sure that you're staying up to date, even just like trying to stay up to date with the news as we were talking about. So I would love to hear your take on on the different maturity levels that you've been seeing and what that looks like.
|
||||
|
||||
Iveta Lohovska:
|
||||
So I have huge exposure to GenAI for the enterprise, and there's a huge component expectation management. Why? Because chat GPT for conversational purposes and individual help is something very cool. But when this needs to be translated into actual business use cases scenario with all the constraint of the enterprise architecture, with the constraint of the use cases, the reality changes quite dramatically. So end users who are used to expect level of forgiveness as conversational chatbots have, is very different of what you will get into actual, let's say, knowledge worker type of context, or summarization type of context into the enterprise. And it's not so much to the performance of the models, but we have something called modalities of the models. And I don't think there will be ultimately one model with all the capabilities possible, let's say cult generation or image generation, voice generational, or just being very chatty and loving and so on. There will be multiple mini models out there for those. Modalities in actual architecture with reasonable cost are very difficult to handle.
|
||||
|
||||
Iveta Lohovska:
|
||||
So I would say the technical community feels we are very mature and very fast. The enterprise adoption is a totally different topic, and it's a couple of years behind, but also the society type of technologists like me, who try to keep up with the development and we know where we stand at this point, but they're the legal side and the regulations coming in, like the EU act and Biden trying to regulate the compute power, but also how societies react to this and how they adapt. And I think especially on the third one, we are far behind understanding and the implications of this technology, also adopting it at scale and understanding the vulnerabilities. That's why I enjoy so much my enterprise work is because it's a reality check. When you put the price tag attached to actual Gen AI use case in production with the inference cost and the expected performance, it's different situation when you just have an app on the phone and you chat with it and it pulls you interesting links. So yes, I think that there's a bridge to be built between the two worlds.
|
||||
|
||||
Demetrios:
|
||||
Yeah. And I find it really interesting too, because it feels to me like since it is so new, people are more willing to explore and not necessarily have that instant return of the ROI, but when it comes to more traditional ML or predictive ML, it is a bit more mature and so there's less patience for that type of exploration. Or, hey, is this use case? If you can't by now show the ROI of a predictive ML use case, then that's a little bit more dangerous. But if you can't with a Gen AI use case, it is not that big of a deal.
|
||||
|
||||
Iveta Lohovska:
|
||||
Yeah, it's basically a technology growing up in front of our eyes. It's a kind of a flying a plane while building it type of situation. We are seeing it in the real time, and I agree with you. So that the maturity around ML is one thing, but around generative AI, and they will be a model of kind of mini disappointment or decline, in my opinion, before actually maturing product. This kind of powerful technology in a sustainable way. Sustainable ways mean you can afford it, but also it proves your business case and use case. Otherwise it's just doing for the sake of doing it because everyone else is doing it.
|
||||
|
||||
Demetrios:
|
||||
Yeah, yeah, 100%. So I know we're bumping up against time here. I do feel like there was a bit of a topic that we wanted to discuss with the licenses and how that plays into basically trustworthiness and explainability. And so we were talking about how, yeah, the best is to run your own model, and it probably isn't going to be this gigantic model that can do everything. It's the, it seems like the trends are going into smaller models. And from your point of view though, we are getting new models like every week. It feels like. Yeah, especially.
|
||||
|
||||
Demetrios:
|
||||
I mean, we were just talking about this before we went live again, like databricks just released there. What is it? DBRX Yesterday you had Mistral releasing like a new base model over the weekend, and then Llama 3 is probably going to come out in the flash of an eye. So where do you stand in regards to that? It feels like there's a lot of movement in open source, but it is a little bit of, as you mentioned, like, to be cautious with the open source movement.
|
||||
|
||||
Iveta Lohovska:
|
||||
So I think it feels like there's a lot of open source, but that. So I'm totally for open sourcing and giving the people and the communities the power to be able to innovate, to do R & D in different labs so it's not locked to the view. Elite big tech companies that can afford this kind of technology. So kudos to meta for trying compared to the other equal players in the space. But open source comes with a lot of ecosystem in our world, especially for the more powerful models, which is something I don't like because it becomes like just, it immediately translates into legal fees type of conversation. It's like there are too many if else statements in those open source licensing terms where it becomes difficult to navigate, for technologists to understand what exactly this means, and then you have to bring the legal people to articulate it to you or to put additional clauses. So it's becoming a very complex environment to handle and less and less open, because there are not so many open source and small startup players that can afford to train foundational models that are powerful and useful. So it becomes a bit of a game logged to a view, and I think everyone needs to be a bit worried about that.
|
||||
|
||||
Iveta Lohovska:
|
||||
So we can use the equivalents from the past, but I don't think we are doing well enough in terms of open sourcing. The three main core components of LLM model, which is the model itself, the data it has been trained on, and the data sets, and most of the times, at least in one of those, is restricted or missing. So it's difficult space to navigate.
|
||||
|
||||
Demetrios:
|
||||
Yeah, yeah. You can't really call it trustworthy, or you can't really get the information that you need and that you would hope for if you're missing one of those three. I do like that little triangle of the necessities. So, Iveta, this has been awesome. I really appreciate you coming on here. Thank you, Sabrina, for joining us. And for everyone else that is watching, remember, don't get lost in vector space. This has been another vector space talk.
|
||||
|
||||
Demetrios:
|
||||
We are out. Have a great weekend, everyone.
|
||||
|
||||
Iveta Lohovska:
|
||||
Thank you. Bye. Thank you. Bye.
|
||||
@@ -1,15 +1,15 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: How to meow on the long tail with Cheshire Cat AI? - Piero and Nicola |
|
||||
Vector Space Talks
|
||||
slug: meow-with-cheshire-cat
|
||||
short_description: Piero Savastano and Nicola Procopio discuss on the ins and
|
||||
short_description: Piero Savastano and Nicola Procopio discusses the ins and
|
||||
outs of Cheshire Cat AI.
|
||||
description: Cheshire Cat AI's Piero Savastano and Nicola Procopio discusses the
|
||||
framework's vector space complexities, community growth, and future
|
||||
cloud-based expansions.
|
||||
preview_image: /blog/from_cms/piero-and-nicola-bp-cropped.png
|
||||
date: 2024-03-18T11:00:13.338Z
|
||||
date: 2024-04-09T03:05:00.000Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -19,7 +19,7 @@ tags:
|
||||
- Vector Search
|
||||
- Vector database
|
||||
---
|
||||
> *"Yes, we love Qdrant. It is our default DB. We support it in three different forms, file based, container based, and cloud based also.”*\
|
||||
> *"We love Qdrant! It is our default DB. We support it in three different forms, file based, container based, and cloud based as well.”*\
|
||||
— Piero Savastano
|
||||
>
|
||||
|
||||
@@ -31,11 +31,11 @@ Piero Savastano is the Founder and Maintainer of the open-source project, Cheshi
|
||||
|
||||
Nicola Procopio has more than 10 years of experience in data science and has worked in different sectors and markets from Telco to Healthcare. At the moment he works in the Media market, specifically on semantic search, vector spaces, and LLM applications. He has worked in the R&D area on data science projects and he has been and is currently a contributor to some open-source projects like Cheshire Cat. He is the author of popular science articles about data science on specialized blogs.
|
||||
|
||||
***Listen to the episode on Spotify, Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on YouTube.***
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/2d58Xui99QaUyXclIE1uuH?si=68c5f1ae6073472f), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/K40DIG9ZzAU?feature=shared).***
|
||||
|
||||
[embed YouTube video here]
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/K40DIG9ZzAU?si=rK0EVXmvNJ5OSZa4" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
[embed anchor.fm podcast here]
|
||||
<iframe src="https://podcasters.spotify.com/pod/show/qdrant-vector-space-talk/embed/episodes/How-to-meow-on-the-long-tail-with-Cheshire-Cat-AI----Piero-and-Nicola--Vector-Space-Talks-018-e2h7k59/a-ab31teu" height="102px" width="400px" frameborder="0" scrolling="no"></iframe>
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
@@ -45,11 +45,11 @@ It’s time to learn how to meow! Piero in this episode of Vector Space Talks di
|
||||
|
||||
Here are the highlights from this episode:
|
||||
|
||||
1. The Art of Embedding: Discover how Cheshire Cat uses collections with an embedder, fine-tuning them through scalar quantization and other methods to enhance accuracy and performance.
|
||||
2. Vectors in Harmony: Get the lowdown on storing quantized vectors in a hybrid mode – it's all about saving memory without compromising on speed.
|
||||
3. Memory Matters: Scoop on managing different types of memory within Qdrant, the go-to vector DB for Cheshire Cat.
|
||||
4. Community Chronicles: Talking about the growing community that's shaping the evolution of Cheshire Cat - from enthusiasts to core contributors!
|
||||
5. Looking Ahead: They've got grand plans brewing for a cloud version of Cheshire Cat. Imagine a marketplace buzzing with user-generated plugins. This is the future they're painting!
|
||||
1. **The Art of Embedding:** Discover how Cheshire Cat uses collections with an embedder, fine-tuning them through scalar quantization and other methods to enhance accuracy and performance.
|
||||
2. **Vectors in Harmony:** Get the lowdown on storing quantized vectors in a hybrid mode – it's all about saving memory without compromising on speed.
|
||||
3. **Memory Matters:** Scoop on managing different types of memory within Qdrant, the go-to vector DB for Cheshire Cat.
|
||||
4. **Community Chronicles:** Talking about the growing community that's shaping the evolution of Cheshire Cat - from enthusiasts to core contributors!
|
||||
5. **Looking Ahead:** They've got grand plans brewing for a cloud version of Cheshire Cat. Imagine a marketplace buzzing with user-generated plugins. This is the future they're painting!
|
||||
|
||||
> Fun Fact: The Cheshire Cat community on Discord plays a crucial role in the development and user support of the framework, described humorously by Piero as "a mess" due to its large and active nature.
|
||||
>
|
||||
@@ -109,10 +109,10 @@ Piero Savastano:
|
||||
Dark team, you can do a lot of stuff with the framework. This is how it presents itself. We have a blog with tutorials, but going back to our numbers, it is open source, GPL licensed. We have some good numbers. We are mostly active in Italy and in a good part of Europe, East Europe, and also a little bit of our communities in the United States. There are a lot of contributors already and our docker image has been downloaded quite a few times, so it's really easy to start up and running because you just docker run and you're good to go. We have also a discord server with thousands of members. If you want to join us, it's going to be fun.
|
||||
|
||||
Piero Savastano:
|
||||
We like meme, we like to build culture around code, so it is not just the code, these are the main components of the cat. You have a chat as usual. The rabbitol is our module dedicated to document ingestion. You can extend all of these parts. We have an agent manager. Meddetter is the module to manage plugins. We have a vectordb which is Qdrant natively, by the way. We use both the file based Qdrant, the container version, and also we support the cloud version.
|
||||
We like meme, we like to build culture around code, so it is not just the code, these are the main components of the cat. You have a chat as usual. The rabbit hole is our module dedicated to document ingestion. You can extend all of these parts. We have an agent manager. Meddetter is the module to manage plugins. We have a vectordb which is Qdrant natively, by the way. We use both the file based Qdrant, the container version, and also we support the cloud version.
|
||||
|
||||
Piero Savastano:
|
||||
So if you are using Qdrant, we support the whole stack. Right now with the framework we have an embedder and a large language model coming to the embedder and language models. You can use any language model or embedded you want, closed source API, open Olama, self hosted anything. These are the main features. So the first feature of the cat is that he's ready to fight. It is already dogrized. It's model agnostic. One command in the terminal and you can meow.
|
||||
So if you are using Qdrant, we support the whole stack. Right now with the framework we have an embedder and a large language model coming to the embedder and language models. You can use any language model or embedded you want, closed source API, open Ollama, self hosted anything. These are the main features. So the first feature of the cat is that he's ready to fight. It is already dogsized. It's model agnostic. One command in the terminal and you can meow.
|
||||
|
||||
Piero Savastano:
|
||||
The other aspect is that there is not only a retrieval augmented generation system, but there is also an action agent. This is all customizable. You can plug in any script you want as an agent, or you can customize the ready default presence default agent. And one of our specialty is that we do retrieve augmented generation, not only on documents as everybody's doing, but we do also augmented generation over conversations. I can hear your keyboard. We do augmented generation over conversations and over procedures. So also our tools and form conversational forms are embedded into the DB. We have a big plugin system.
|
||||
@@ -145,10 +145,10 @@ Nicola Procopio:
|
||||
This collection with the name of the embedder used. When the user changed the embedder, we check if the embedder has the same dimension. If has the same dimension, we check also the aliases. If the aliases is the same we don't change nothing. Otherwise we create another collection and this is the drunken cut effect. The first feature that we use in the cat. Another feature is the quantization because with this Qdrant feature we improve the accuracy at the performance. We use the scalar quantitation because we are model agnostic and other quantitation like the binary quantitation.
|
||||
|
||||
Nicola Procopio:
|
||||
If you read on the Qdrant documents are experimented on not to all embedder but also for OpenAI and Coer. If I remember well with this discover quantitation and the scour quantization is used in the storage step. The vector are quantizzed and stored in a hybrid mode, the original vector on disk, the quantized vector in RAm and with this procedure we procedure we can use less memory. In case of Qdrant scalar quantization, the flat 32 elements is converted to int eight on a single number on a single element needs 75% less memory. In case of big embeddings like I don't know Gina embeddings or mistral embeddings with more than 1000 elements. This is big improvements. The second part is the retriever step. We use a quantizement query at the quantized vector to calculate causing similarity and we have the top n results like a simple semantic search pipeline.
|
||||
If you read on the Qdrant documents are experimented on not to all embedder but also for OpenAI and Coer. If I remember well with this discover quantitation and the scour quantization is used in the storage step. The vector are quantized and stored in a hybrid mode, the original vector on disk, the quantized vector in RAM and with this procedure we procedure we can use less memory. In case of Qdrant scalar quantization, the flat 32 elements is converted to int eight on a single number on a single element needs 75% less memory. In case of big embeddings like I don't know Gina embeddings or mistral embeddings with more than 1000 elements. This is big improvements. The second part is the retriever step. We use a quantizement query at the quantized vector to calculate causing similarity and we have the top n results like a simple semantic search pipeline.
|
||||
|
||||
Nicola Procopio:
|
||||
But if we want a top end results in quantiz mod, the quantity mod has less quality on the information and we use the oversampling. The oversampling is a simple multiplication. If we want top n with n ten with oversampling with a score like one five, we have 15 results, quantities results. When we have these 15 quantities results, we retrieve also the same 15 unquanted vectors. And on these unquanted vectors we rescale busset on the query and filter the best ten. This is an improvement because the retrieve step is so fast. Yes, because using these tip and tricks, the cheshire capped vectors achieve up.
|
||||
But if we want a top end results in quantize mod, the quantity mod has less quality on the information and we use the oversampling. The oversampling is a simple multiplication. If we want top n with n ten with oversampling with a score like one five, we have 15 results, quantities results. When we have these 15 quantities results, we retrieve also the same 15 unquanted vectors. And on these unquanted vectors we rescale busset on the query and filter the best ten. This is an improvement because the retrieve step is so fast. Yes, because using these tip and tricks, the Cheshire capped vectors achieve up.
|
||||
|
||||
Piero Savastano:
|
||||
Four.
|
||||
@@ -172,7 +172,7 @@ Demetrios:
|
||||
Then we do that, we can get the full program. How cool is that? Well, let's see, I'll give it another minute, let anyone from the chat ask any questions. This was really cool and I appreciate you all breaking down. Not only the space and what you're doing, but the different ways that you're using Qdrant and the challenges and the architecture behind it. I would love to know while people are typing in their questions, especially for you, Nicola, what have been some of the challenges that you've faced when you're dealing with just trying to get Cheshire Cat to be more reliable and be more able to execute with confidence?
|
||||
|
||||
Nicola Procopio:
|
||||
The challenges are in particular to mix a lot of Qdrant feature with the user needs. Because I'm a researcher, a data scientist, I like to play with strange features like binary quantization, but we need to maintain the focus on the user needs, on the user behavior. And sometimes we cut some feature on the Chichircat because it's not important now for for the user and we can introduce some bug, or rather misunderstanding for the user.
|
||||
The challenges are in particular to mix a lot of Qdrant feature with the user needs. Because I'm a researcher, a data scientist, I like to play with strange features like binary quantization, but we need to maintain the focus on the user needs, on the user behavior. And sometimes we cut some feature on the Cheshire cat because it's not important now for for the user and we can introduce some bug, or rather misunderstanding for the user.
|
||||
|
||||
Demetrios:
|
||||
Can you hear me? Yeah. All right, good. Now I'm seeing a question come through in the chat that is asking if you are thinking about cloud version of the cat. Like a SaaS, it's going to come. It's in the works.
|
||||
@@ -190,7 +190,7 @@ Demetrios:
|
||||
Yeah, that's the best. That is really cool. Simone is asking if there's companies that are already using Cheshire cat, and if you can mention a few.
|
||||
|
||||
Piero Savastano:
|
||||
Yeah, okay. In Italy, there are at least 1015 companies distributed along education, customer care, typical chatbot usage. Also, one of them in particular is trying to build for public administration, which is really hard to do on the international level. We are seeing something in Germany, like web agencies starting to use the cat a little on the USA. Mostly they are trying to build agents using the cat and Olama as a runner. And a company in particular presented in a conference in Vegas a pitch about a 3d avatar. Inside the avatar, there is the cat as a linguistic device.
|
||||
Yeah, okay. In Italy, there are at least 1015 companies distributed along education, customer care, typical chatbot usage. Also, one of them in particular is trying to build for public administration, which is really hard to do on the international level. We are seeing something in Germany, like web agencies starting to use the cat a little on the USA. Mostly they are trying to build agents using the cat and Ollama as a runner. And a company in particular presented in a conference in Vegas a pitch about a 3d avatar. Inside the avatar, there is the cat as a linguistic device.
|
||||
|
||||
Demetrios:
|
||||
Oh, nice.
|
||||
@@ -199,7 +199,7 @@ Piero Savastano:
|
||||
To be honest, we have a little problem tracking companies because we still have no telemetry. We decided to be no telemetry for the moment. So I hope companies will contribute and make themselves happen. If that does not, we're going to track a little more. But companies using the cat are at least in the 50, 60, 70.
|
||||
|
||||
Demetrios:
|
||||
Yeah, nice. So if anybody out there is using the cat, and you have not talked to Piro yet, let him know so that he can have a good idea of what you're doing and how you're doing it. There's also another question coming through about the market analysis. Are there some competitors?
|
||||
Yeah, nice. So if anybody out there is using the cat, and you have not talked to Piero yet, let him know so that he can have a good idea of what you're doing and how you're doing it. There's also another question coming through about the market analysis. Are there some competitors?
|
||||
|
||||
Piero Savastano:
|
||||
There are many competitors. When you go down to what distinguishes the cat from many other frameworks that are coming out, we decided since the beginning to go for a plugin based operational agent. And at the moment, most frameworks are retrieval augmented generation frameworks. We have both retrieval augmented generation. We have tooling, we have forms. The tools and the forms are also embedded. So the cat can have 20,000 tools, because we also embed the tools and we make a recall over the function calling. So we scaled up both documents, conversation and tools, conversational forms, and I've not seen anybody doing that till now.
|
||||
|
||||
@@ -1,16 +1,16 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: Insight Generation Platform for LifeScience Corporation - Hooman
|
||||
Sedghamiz | Vector Space Talks
|
||||
slug: insight-generation-platform
|
||||
short_description: Hooman Sedghamiz explores the untapped potential of large
|
||||
language models in creating cutting-edge applications.
|
||||
description: Hooman Sedghamiz unfolds the potential of AI in life sciences, from
|
||||
custom knowledge applications to improving crop yield predictions, while
|
||||
teasing apart the nuances of in-house AI deployment for multi-faceted
|
||||
short_description: Hooman Sedghamiz explores the potential of large language
|
||||
models in creating cutting-edge AI applications.
|
||||
description: Hooman Sedghamiz discloses the potential of AI in life sciences,
|
||||
from custom knowledge applications to improving crop yield predictions, while
|
||||
tearing apart the nuances of in-house AI deployment for multi-faceted
|
||||
enterprise efficiency.
|
||||
preview_image: /blog/from_cms/hooman-sedghamiz-bp-cropped.png
|
||||
date: 2024-03-08T09:45:59.753Z
|
||||
date: 2024-03-25T08:46:28.227Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -18,11 +18,13 @@ tags:
|
||||
- Retrieval Augmented Generation
|
||||
- Insight Generation Platform
|
||||
---
|
||||
> *"So there is this really great vector db comparison that came out recently. I saw there are like maybe more than 40 vector stores in 2024. When we started back in 2023 was only a few. And what I see, which is really lacking in this pipeline of retrieval augmented generation is major innovation around data pipeline.”*\
|
||||
> *"There is this really great vector db comparison that came out recently. I saw there are like maybe more than 40 vector stores in 2024. When we started back in 2023, there were only a few. What I see, which is really lacking in this pipeline of retrieval augmented generation is major innovation around data pipeline.”*\
|
||||
-- Hooman Sedghamiz
|
||||
>
|
||||
|
||||
Hooman Sedghamiz**,** Sr. Director AI/ML - Insights at Bayer AG is a distinguished figure in AI and ML in the life sciences field. With a wealth of experience, he has led teams and projects that have greatly advanced medical products, including implantable and wearable devices. Notably, he served as the Generative AI product owner and senior director at Bayer Pharmaceuticals, where he played a pivotal role in developing a GPT-based central platform for precision medicine. In 2023, he assumed the role of Co-Chair for the EMNLP 2023 GEM industrial track, furthering his contributions to the field. Hooman has also been an AI/ML advisor and scientist at the University of California, San Diego, leveraging his expertise in deep learning to drive biomedical research and innovation. His strengths lie in guiding data science initiatives from inception to commercialization and bridging the gap between medical and healthcare applications through MLOps, LlmOps, and deep learning product management. Engaging with research institutions and collaborating closely with Dr. Nemati at Harvard University and UCSD, Hooman continues to be a dynamic and influential figure in the data science community.
|
||||
Hooman Sedghamiz, Sr. Director AI/ML - Insights at Bayer AG is a distinguished figure in AI and ML in the life sciences field. With years of experience, he has led teams and projects that have greatly advanced medical products, including implantable and wearable devices. Notably, he served as the Generative AI product owner and Senior Director at Bayer Pharmaceuticals, where he played a pivotal role in developing a GPT-based central platform for precision medicine.
|
||||
|
||||
In 2023, he assumed the role of Co-Chair for the EMNLP 2023 GEM industrial track, furthering his contributions to the field. Hooman has also been an AI/ML advisor and scientist at the University of California, San Diego, leveraging his expertise in deep learning to drive biomedical research and innovation. His strengths lie in guiding data science initiatives from inception to commercialization and bridging the gap between medical and healthcare applications through MLOps, LLMOps, and deep learning product management. Engaging with research institutions and collaborating closely with Dr. Nemati at Harvard University and UCSD, Hooman continues to be a dynamic and influential figure in the data science community.
|
||||
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/2oj2ne5l9qrURQSV0T1Hft?si=DMJRTAt7QXibWiQ9CEKTJw), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/yfzLaH5SFX0).***
|
||||
|
||||
@@ -32,9 +34,9 @@ Hooman Sedghamiz**,** Sr. Director AI/ML - Insights at Bayer AG is a distinguish
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
Why is real-time evaluation critical in maintaining the integrity of chatbot interactions and preventing issues like promoting competitors or making false promises? What strategies do developers employ to minimize cost while maximizing the effectiveness of model evaluations, specifically when dealing with LLMs? These might be just some of the many questions people in the industry are asking themselves. Worry not! Because Demetrios and Sourabh will break it down for you.
|
||||
Why is real-time evaluation critical in maintaining the integrity of chatbot interactions and preventing issues like promoting competitors or making false promises? What strategies do developers employ to minimize cost while maximizing the effectiveness of model evaluations, specifically when dealing with LLMs? These might be just some of the many questions people in the industry are asking themselves. We aim to cover most of it in this talk.
|
||||
|
||||
Check out their conversation as they dive into the intricate world of AI chatbot evaluations. Discover the nuances of ensuring your chatbot's quality and continuous improvement across various metrics.
|
||||
Check out their conversation as they peek into world of AI chatbot evaluations. Discover the nuances of ensuring your chatbot's quality and continuous improvement across various metrics.
|
||||
|
||||
Here are the key topics of this episode:
|
||||
|
||||
@@ -44,7 +46,8 @@ Here are the key topics of this episode:
|
||||
4. **Cost-Effective Evaluation Models**: Discussion on employing smaller models for evaluation to reduce costs without compromising the depth of analysis, focusing on failure cases and root-cause assessments.
|
||||
5. **Tailored Evaluation Metrics**: Emphasis on the necessity of customizing evaluation criteria to suit specific use case requirements, including an exploration of the different metrics applicable to diverse scenarios.
|
||||
|
||||
> Fun Fact: Sourabh discussed the use of Uptrend, an innovative API that provides scores and explanations for various data checks, facilitating logical and informed decision-making when evaluating AI models.
|
||||
>Fun Fact: Large language models like Mistral, Llama, and Nexus Raven have improved in their ability to perform function calling with low hallucination and high-quality output.
|
||||
|
||||
>
|
||||
|
||||
## Show notes:
|
||||
@@ -73,13 +76,16 @@ Here are the key topics of this episode:
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
We are here and I couldn't think of a better way to spend my Valentine's Day than with you humen this is absolutely incredible. I'm so excited for this talk that you're going to bring and I want to let everyone that is out there listening know what caliber of of speaker we have with us today because you have done a lot of stuff. Folks out there do not let this man's young look fool you. You look like you are not in your. When it comes to your bio, it looks like you should be in your very excited. You've got a lot of experience running data science projects, ML projects, LLM projects, all that fun stuff. You're working at Bayern Munich, sorry, not Bayern Munich, Bayer AG. And you're the senior director of AI and ML.
|
||||
We are here and I couldn't think of a better way to spend my Valentine's Day than with you Hooman this is absolutely incredible. I'm so excited for this talk that you're going to bring and I want to let everyone that is out there listening know what caliber of a speaker we have with us today because you have done a lot of stuff. Folks out there do not let this man's young look fool you. You look like you are not in your fifty's or sixty's. But when it comes to your bio, it looks like you should be in your seventy's. I am very excited. You've got a lot of experience running data science projects, ML projects, LLM projects, all that fun stuff. You're working at Bayern Munich, sorry, not Bayern Munich, Bayer AG. And you're the senior director of AI and ML.
|
||||
|
||||
|
||||
Demetrios:
|
||||
And I think that there is a ton of other stuff that you've done when it comes to machine learning, artificial intelligence. You've got both like the traditional ML background, I think, and then you've also got this new generative AI background and so you can leverage both. But you also think about things in data engineering way. You understand the whole lifecycle. And so today we get to talk all about some of this fun. I know you've got some slides prepared for us. I'll let you throw those on and I'll let anyone else in the chat. Feel free to ask questions while humin is going through the presentation and I'll jump in and stop them when needed.
|
||||
And I think that there is a ton of other stuff that you've done when it comes to machine learning, artificial intelligence. You've got both like the traditional ML background, I think, and then you've also got this new generative AI background and so you can leverage both. But you also think about things in data engineering way. You understand the whole lifecycle. And so today we get to talk all about some of this fun. I know you've got some slides prepared for us. I'll let you throw those on and I'll let anyone else in the chat. Feel free to ask questions while Hooman is going through the presentation and I'll jump in and stop them when needed.
|
||||
|
||||
|
||||
Demetrios:
|
||||
But also we can have a little discussion after a few minutes of slides. So for everyone looking, we're going to be watching this and then we're going to be checking out like really talking about what 2024 AI in the enterprise looks like and what is needed to really take advantage of that. So human, I'm dropping off to you, man, and I'll jump in when needed.
|
||||
But also we can have a little discussion after a few minutes of slides. So for everyone looking, we're going to be watching this and then we're going to be checking out like really talking about what 2024 AI in the enterprise looks like and what is needed to really take advantage of that. So Hooman, I'm dropping off to you, man, and I'll jump in when needed.
|
||||
|
||||
|
||||
Hooman Sedghamiz:
|
||||
Thanks a lot for the introduction. Let me get started. Do you have my screen already?
|
||||
@@ -94,10 +100,10 @@ Hooman Sedghamiz:
|
||||
So now you can imagine via is really important to us because it has the potential of unlocking a future where good health is a reality and hunger is a memory. So I maybe start about maybe giving you a hint of what are really the numerous use cases that AI or challenges that AI could help out with. In life science industry. You can think of adverse event detection when patients are taking a medication, too much of it. The patients might report adverse events, stomach bleeding and go to social media post about it. A few years back, it was really difficult to process automatically all this sort of natural text in a kind of scalable manner. But nowadays, thanks to large language models, it's possible to automate this and identify if there is a medication or anything that might have negatively an adverse event on a patient population. Similarly, you can now create a lot of marketing content using these large language models for products.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
At the same time, drug discovery is making really big strides when it comes to identifying new compounds. You can essentially describe these compounds using formats like smiles, which could be represented as really text. And these large language models can be trained on them and they can predict the sequences. At the same time, you have this clinical trial outcome prediction, which is huge for pharmaceutical companies. If you could predict what will be the outcome of a trial, it would be a huge time and resource saving for a lot of companies. And of course, a lot of us already see in the market a lot of medical virtual assistants using large language models that can answer medical inquiries and give consultations around them. And there is really, I believe the biggest potential here is around real world data, like most of us nowadays, have some sort of sensor or watch that's measuring our health maybe at a minute by minute level, or it's measuring our heart rate. You go to the hospital, you have all your medical records recorded there, and these large language models have their capacity to process this complex data, and you will be able to drive better insights for individualized insights for patients.
|
||||
At the same time, drug discovery is making really big strides when it comes to identifying new compounds. You can essentially describe these compounds using formats like smiles, which could be represented as real text. And these large language models can be trained on them and they can predict the sequences. At the same time, you have this clinical trial outcome prediction, which is huge for pharmaceutical companies. If you could predict what will be the outcome of a trial, it would be a huge time and resource saving for a lot of companies. And of course, a lot of us already see in the market a lot of medical virtual assistants using large language models that can answer medical inquiries and give consultations around them. And there is really, I believe the biggest potential here is around real world data, like most of us nowadays, have some sort of sensor or watch that's measuring our health maybe at a minute by minute level, or it's measuring our heart rate. You go to the hospital, you have all your medical records recorded there, and these large language models have their capacity to process this complex data, and you will be able to drive better insights for individualized insights for patients.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
And our company is also in crop science, as I mentioned, and crop yield prediction. If you could help farmers improve their crop yield, it means that they can produce better products faster with higher quality. So maybe I could start with maybe a history in 2023, what happened? How companies like ours were looking at large language models and opportunities. They bring, I think in 2023, everyone was excited to bring these efficiency games, right? Everyone wanted to use them for creating content, drafting emails, all these really low hanging fruit use cases. That was around. And one of the earlier really nice architectures that came up that I really like was from a 16 z enterprise that was, I think, back in really, really early 2023. LangChain was new, we had land chain and we had all this. Of course, Kudran been there for a long time, but it was the first time that you could see vector store products could be integrated into applications.
|
||||
And our company is also in crop science, as I mentioned, and crop yield prediction. If you could help farmers improve their crop yield, it means that they can produce better products faster with higher quality. So maybe I could start with maybe a history in 2023, what happened? How companies like ours were looking at large language models and opportunities. They bring, I think in 2023, everyone was excited to bring these efficiency games, right? Everyone wanted to use them for creating content, drafting emails, all these really low hanging fruit use cases. That was around. And one of the earlier really nice architectures that came up that I really like was from a 16 z enterprise that was, I think, back in really, really early 2023. LangChain was new, we had land chain and we had all this. Of course, Qdrant been there for a long time, but it was the first time that you could see vector store products could be integrated into applications.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
Really at large scale. There are different components. It's quite complex architecture. So on the right side you see how you can host large language models. On the top you see how you can augment them using external data. Of course, we had these plugins, right? So you can connect these large language models with Google search APIs, all those sort of things, and some validation that are in the middle that you could use to validate the responses fast forward. Maybe I can kind of spend, let me check out the time. Maybe I can spend a few minutes about the components of LLM APIs and hosting because that I think has a lot of potential in terms of applications that need to be really scalable.
|
||||
@@ -109,10 +115,10 @@ Hooman Sedghamiz:
|
||||
And we know that most of the users won't use that. We know that it's a usage based application. You just probably go there. Depending on your daily work, you probably use it. Some people don't use it heavily. I kind of did some calculation. If you build it in house using APIs that you can access yourself, and large language models that corporations can deploy internally and locally, that cost saving could be huge, really magnitudes cheaper, maybe 30 to 20 to 30 times cheaper. So looking, comparing 2024 to 2023, a lot of things have changed.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
Like if you look at the open source large language models that came out really great models from Mistral, now we have models like Llama, two based model, all of these models came out. You can now kind of take a look and see that the performance of them is really, really getting close, if not better than GPT 3.5 already at same level and really approaching step by step to GPT four. And looking at the price on the right side and speed or throughput, you can see that like for example, Mixtrawl seven eight B could be a really cheap option to deploy. And also the performance of it gets really close to GPT 3.5 for many use cases in the enterprise companies. I think two of the big things this year, end of last year that came out that make this kind of really a reality are really a few large language models. I don't know if I can call them large language models. They are like 7 billion to 13 billion compared to GPT four, GT 3.5. I don't think they are really large.
|
||||
Like if you look at the open source large language models that came out really great models from Mistral, now we have models like Llama, two based model, all of these models came out. You can now kind of take a look and see that the performance of them is really, really getting close, if not better than GPT 3.5 already at same level and really approaching step by step to GPT 4. And looking at the price on the right side and speed or throughput, you can see that like for example, Mistral seven eight B could be a really cheap option to deploy. And also the performance of it gets really close to GPT 3.5 for many use cases in the enterprise companies. I think two of the big things this year, end of last year that came out that make this kind of really a reality are really a few large language models. I don't know if I can call them large language models. They are like 7 billion to 13 billion compared to GPT four, GT 3.5. I don't think they are really large.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
But one was Nexus, Raven. We know that applications, if they want to be robust, they really need function calling. We are seeing this paradigm of function calling, which essentially you ask a language model to generate structured output, you give it a function signature, right? You ask it to generate an output, structured output argument for that function. Next was Raven came out last year, that, as you can see here, really is getting really close to GPT four, right? And GPT four being magnitude bigger than this model. This model only being 13 billion parameters really provides really less hallucination, but at the same time really high quality of function calling. So this makes me really excited for the open source and also the companies that want to build their own applications that requires function calling. That was really lacking maybe just five months ago. At the same time, we have really dedicated large language models to programming languages or scripting like SQL, that we are also seeing like SQL coder that's already beating GPT four.
|
||||
But one was Nexus Raven. We know that applications, if they want to be robust, they really need function calling. We are seeing this paradigm of function calling, which essentially you ask a language model to generate structured output, you give it a function signature, right? You ask it to generate an output, structured output argument for that function. Next was Raven came out last year, that, as you can see here, really is getting really close to GPT four, right? And GPT four being magnitude bigger than this model. This model only being 13 billion parameters really provides really less hallucination, but at the same time really high quality of function calling. So this makes me really excited for the open source and also the companies that want to build their own applications that requires function calling. That was really lacking maybe just five months ago. At the same time, we have really dedicated large language models to programming languages or scripting like SQL, that we are also seeing like SQL coder that's already beating GPT four.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
So maybe we can now quickly take a look at how model solving will look like for a large company like ours, like companies that have a lot of people across the globe again, in this aspect also, the community has made really big progress, right? So we have text generation inference from hugging face is open source for most purposes, can be used and it's the choice of mine and probably my group prefers this option. But we have Olama, which is great, a lot of people are using it. We have llama CPP which really optimizes the large language models for local deployment as well, and edge devices. I was really amazed seeing Raspberry PI running a large language model, right? Using Llama CPP. And you have this text generation inference that offers quantization support, continuous patching, all those sort of things that make these large LLMs more quantized or more compressed and also more suitable for deployment to large group of people. Maybe I can kind of give you kind of a quick summary of how, if you decide to deploy these large language models, what techniques you could use to make them more efficient, cost friendly and more scalable. So we have a lot of great open source projects like we have Lite LLM which essentially creates an open AI kind of signature on top of your large language models that you have deployed. Let's say you want to use Azure to host or to access GPT four gypty 3.5 or OpenAI to access OpenAI API.
|
||||
|
||||
@@ -1,17 +1,17 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: Production-scale RAG for Real-Time News Distillation - Robert Caulk |
|
||||
Vector Space Talks
|
||||
slug: real-time-news-distillation-rag
|
||||
short_description: Robert Caulk dives into the challenges and innovations in
|
||||
open source AI and news article modeling
|
||||
short_description: Robert Caulk tackles the challenges and innovations in open
|
||||
source AI and news article modeling.
|
||||
description: Robert Caulk, founder of Emergent Methods, discusses the
|
||||
intricacies of context engineering, the power of Newscatcher API for broader
|
||||
complexities of context engineering, the power of Newscatcher API for broader
|
||||
news access, and the sophisticated use of tools like Qdrant for improved
|
||||
recommendation systems, all while emphasizing the importance of efficiency and
|
||||
modularity in technology stacks for real-time data management.
|
||||
preview_image: /blog/from_cms/robert-caulk-bp-cropped.png
|
||||
date: 2024-03-08T10:07:45.278Z
|
||||
date: 2024-03-25T08:49:22.422Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -36,14 +36,14 @@ Robert, Founder of Emergent Methods is a scientist by trade, dedicating his care
|
||||
|
||||
How do Robert Caulk and Emergent Methods contribute to the open-source community, particularly in AI systems and news article modeling?
|
||||
|
||||
In this episode, we're getting under the hood of open-source projects that are reshaping how we interact with AI systems and news article modeling. Robert takes us on a deep dive into the evolving landscape of news distribution and the tech making it more efficient and balanced.
|
||||
In this episode, we'll be learning stuff about open-source projects that are reshaping how we interact with AI systems and news article modeling. Robert takes us on an exploration into the evolving landscape of news distribution and the tech making it more efficient and balanced.
|
||||
|
||||
Here are some takeaways from this episode:
|
||||
|
||||
1. **Context Matters**: Discover the importance of context engineering in news and how it ensures a diversified and consumable information flow.
|
||||
2. **Introducing Newscatcher API**: Get the lowdown on how this tool taps into 50,000 news sources for more thorough and up-to-date reporting.
|
||||
3. **The Magic of Embedding**: Learn about article summarization and semantic search, and how they're crucial for discovering content that truly resonates.
|
||||
4. **Quadrant & Cloud**: Explore how Qdrant's cloud offering and its single responsibility principle support a robust, modular approach to managing news data.
|
||||
4. **Qdrant & Cloud**: Explore how Qdrant's cloud offering and its single responsibility principle support a robust, modular approach to managing news data.
|
||||
5. **Startup Superpowers**: Find out why startups have an edge in implementing new tech solutions and how incumbents are tied down by legacy products.
|
||||
|
||||
> Fun Fact: Did you know that startups' lack of established practices is actually a superpower in the face of new tech paradigms? Legacy products can't keep up!
|
||||
@@ -75,7 +75,7 @@ Here are some takeaways from this episode:
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
Robert, it's great to have you here for the vector space talks. I don't know if you're familiar with some of this fun stuff that we do here, but we get to talk with all kinds of experts like yourself on what they're doing when it comes to the vector space and how you've overcome challenges, how you're working through things, because this is a very new field and it is not the most intuitive, as you will tell us more in this upcoming talk. I really am excited because you've been a scientist by trade. Now, you're currently founder at emergent Methods and you've dedicated your career to a variety of open source projects that range from the large scale AI systems to the discrete element modeling. Now at emergent methods, you are adaptively modeling over 1 million news articles per day. That sounds like a whole lot of news articles. And you've been talking and working through production grade rag, which is basically everyone's favorite topic these days. So I know you got to talk for us, man.
|
||||
Robert, it's great to have you here for the vector space talks. I don't know if you're familiar with some of this fun stuff that we do here, but we get to talk with all kinds of experts like yourself on what they're doing when it comes to the vector space and how you've overcome challenges, how you're working through things, because this is a very new field and it is not the most intuitive, as you will tell us more in this upcoming talk. I really am excited because you've been a scientist by trade. Now, you're currently founder at Emergent Methods and you've dedicated your career to a variety of open source projects that range from the large scale AI systems to the discrete element modeling. Now at emergent methods, you are adaptively modeling over 1 million news articles per day. That sounds like a whole lot of news articles. And you've been talking and working through production grade RAG, which is basically everyone's favorite topic these days. So I know you got to talk for us, man.
|
||||
|
||||
Demetrios:
|
||||
I'm going to hand it over to you. I'll bring up your screen right now, and when someone wants to answer or ask a question, feel free to throw it in the chat and I'll jump out at Robert and stop him if needed.
|
||||
@@ -87,22 +87,22 @@ Demetrios:
|
||||
Great to have you here, man. I'm excited for this one.
|
||||
|
||||
Robert Caulk:
|
||||
Thanks for having me, Demetrius. Yeah, it's a great opportunity. I love talking about vector spaces, parameter spaces. So to talk on the show is great. We've got a lot of fun challenges ahead of us in the industry, I think, and the industry is establishing best practices. Like you said, everybody's just trying to figure out what's going on. And some of these base layer tools like Qdrant really enable products and enable companies and they enable us. So let me start.
|
||||
Thanks for having me, Demetrios. Yeah, it's a great opportunity. I love talking about vector spaces, parameter spaces. So to talk on the show is great. We've got a lot of fun challenges ahead of us in the industry, I think, and the industry is establishing best practices. Like you said, everybody's just trying to figure out what's going on. And some of these base layer tools like Qdrant really enable products and enable companies and they enable us. So let me start.
|
||||
|
||||
Robert Caulk:
|
||||
Yeah, like you said, I'm Robert and I'm a founder of emergent methods. Our background, like you said, we are really committed to free and open source software. We started with a lot of narrow AI. Freak AI was one of our original projects, which is AI ML for algo trading very narrow AI, but we came together and built flowdapt. It's a really nice cluster orchestration software, and I'll talk a little bit about that during this presentation. But some of our background goes into, like you said, large scale deep learning for supercomputers. Really cool, interesting stuff. We have some cloud experience.
|
||||
|
||||
Robert Caulk:
|
||||
We really like configuration, so let's dive into it. Why do we actually need to engineer context in the news? There's a lot of reasons why news is important and why it needs to be distributed in a way that's balanced and diversified, but also consumable. Right, let's look at Chat GPT on the left. This is Chat GPT plus it's kind of hanging out searching for Gaza news on Bing, trying to find the top three articles live. Web search is powerful, but it's slow and ultimately inaccurate. What we're building is real time indexing and we couldn't do that without Qdrant, and there's a lot of reasons which I'll be perfectly happy to dive into, but eventually Chappa Chi PT will pull something together here. There it is. And the first thing it reports is 25 day old article with 25 day old nudes.
|
||||
We really like configuration, so let's dive into it. Why do we actually need to engineer context in the news? There's a lot of reasons why news is important and why it needs to be distributed in a way that's balanced and diversified, but also consumable. Right, let's look at Chat GPT on the left. This is Chat GPT plus it's kind of hanging out searching for Gaza news on Bing, trying to find the top three articles live. Web search is powerful, but it's slow and ultimately inaccurate. What we're building is real time indexing and we couldn't do that without Qdrant, and there's a lot of reasons which I'll be perfectly happy to dive into, but eventually Chat GPT will pull something together here. There it is. And the first thing it reports is 25 day old article with 25 day old nudes.
|
||||
|
||||
Robert Caulk:
|
||||
Old news. So it's just inaccurate. So it's borderline dangerous, what's happening here. Right, so this is a very delicate topic. Engineering context in news properly, which takes a lot of energy, a lot of time and dedication and focus, and not every company really has this sort of resource. So we're talking about enforcing journalistic standards, right? OpenAI and Chachipt, they just don't have the time and energy to build a dedicated prompt for this sort of thing. It's fine, they're doing great stuff, they're helping you code. But someone needs to step in and really do enforce some journalistic standards here.
|
||||
Old news. So it's just inaccurate. So it's borderline dangerous, what's happening here. Right, so this is a very delicate topic. Engineering context in news properly, which takes a lot of energy, a lot of time and dedication and focus, and not every company really has this sort of resource. So we're talking about enforcing journalistic standards, right? OpenAI and Chat GPt, they just don't have the time and energy to build a dedicated prompt for this sort of thing. It's fine, they're doing great stuff, they're helping you code. But someone needs to step in and really do enforce some journalistic standards here.
|
||||
|
||||
Robert Caulk:
|
||||
And that includes enforcing diversity, languages, regions and sources. If I'm going to read about Gaza, what's happening over there, you can bet I want to know what Egypt is saying and what France is saying and what Algeria is saying. So let's do this right. That's kind of what we're suggesting, and the only way to do that is to parse a lot of articles. That's how you avoid outdated, stale reporting. And that's a real danger, which is kind of what we saw on that first slide. Everyone here knows hallucination is a problem and it's something you got to minimize, especially when you're talking about the news. It's just a really high cost if you get it wrong.
|
||||
|
||||
Robert Caulk:
|
||||
And so you need people dedicated to this. And if you're going to dedicate a ton of resources and ton of people, you might as well scale that properly. So that's kind of where this comes into. We call this context engineering news context engineering, to be precise, before llama two, which also is enabling products left and right. As we all know, the traditional pipeline was chunk it up, take 512 tokens, put it through a translator, put it through distillbart, do some sentence extraction, and maybe text classification, if you're lucky, get some sentiment out of it and it works. It gets you something. But after we're talking about reading full articles, getting real rich, context, flexible output, translating, summarizing, really deciding that custom extraction on the fly as your product evolves, that's something that the traditional pipeline really just doesn't support. Right.
|
||||
And so you need people dedicated to this. And if you're going to dedicate a ton of resources and ton of people, you might as well scale that properly. So that's kind of where this comes into. We call this context engineering news context engineering, to be precise, before llama two, which also is enabling products left and right. As we all know, the traditional pipeline was chunk it up, take 512 tokens, put it through a translator, put it through distill art, do some sentence extraction, and maybe text classification, if you're lucky, get some sentiment out of it and it works. It gets you something. But after we're talking about reading full articles, getting real rich, context, flexible output, translating, summarizing, really deciding that custom extraction on the fly as your product evolves, that's something that the traditional pipeline really just doesn't support. Right.
|
||||
|
||||
Robert Caulk:
|
||||
We're talking being able to on the fly say, you know what, actually we want to ask this very particular question of all articles and get this very particular field out. And it's really just a prompt modification. This all is based on having some very high quality, base level, diversified news. And so we'll talk a little bit more. But newscatchers is one of the sources that we're using, which opens up 50,000 different sources. So check them out. That's newscatcherapi.com. They even give free access to researchers if you're doing research in this.
|
||||
@@ -114,10 +114,10 @@ Robert Caulk:
|
||||
And then how do we connect the dots here? Of course, there are many ways to go about it. One way which is interesting and fun to talk about is ide. So that's basically a hypothetical document embedding. And what you do is you use the LLM directly to generate a fake article. And that's what we're showing here on the right. So let's say if the user says, what's going on in New York City government, well, you could say, hey, write me just a hypothetical summary based, it could completely fake and use that to create a fake embedding page and use that for the search. Right. So then you're getting a lot closer to where you want to go.
|
||||
|
||||
Robert Caulk:
|
||||
There's some limitations to this, to it's, there's a computational cost also, it's not updated. It's based on whatever. It's basically diving into what it knows about the New York City government and just creating keywords for you. So there's definitely optimizations here as well. When you talk about ambiguity, well, what if the user follows up and says, well, why did they change the rules? Of course, that's where you can start prompt engineering a little bit more and saying, okay, given this historic conversation and the current question, give me some explicit question without ambiguity, and then do the high de, if that's something you want to do. The real goal here is to stay in a single parameter space, a single vector space. Stay as close as possible when you're doing your search as when you do your embedding. So we're talking here about production scale of stuff.
|
||||
There's some limitations to this, to it's, there's a computational cost also, it's not updated. It's based on whatever. It's basically diving into what it knows about the New York City government and just creating keywords for you. So there's definitely optimizations here as well. When you talk about ambiguity, well, what if the user follows up and says, well, why did they change the rules? Of course, that's where you can start prompt engineering a little bit more and saying, okay, given this historic conversation and the current question, give me some explicit question without ambiguity, and then do the high, if that's something you want to do. The real goal here is to stay in a single parameter space, a single vector space. Stay as close as possible when you're doing your search as when you do your embedding. So we're talking here about production scale of stuff.
|
||||
|
||||
Robert Caulk:
|
||||
So I really am happy to geek out about the stack, the open source stack that we're relying on, which includes Qdrant here. But let's start with Vllm. I don't know if you guys have heard of it. This is a really great new project, and their focus on continuous batching and page detention. And if I'm being completely honest with you, it's really above my pay grade in the technicals and how they're actually implementing all of that inside the GPU memory. But what we do is we outsource that to that project and we really like what they're doing, and we've seen really good results. It's increasing throughput. So when you're talking about trying to parse through a million articles, you're going to need a lot of throughput.
|
||||
So I really am happy to geek out about the stack, the open source stack that we're relying on, which includes Qdrant here. But let's start with VLLM. I don't know if you guys have heard of it. This is a really great new project, and their focus on continuous batching and page detention. And if I'm being completely honest with you, it's really above my pay grade in the technicals and how they're actually implementing all of that inside the GPU memory. But what we do is we outsource that to that project and we really like what they're doing, and we've seen really good results. It's increasing throughput. So when you're talking about trying to parse through a million articles, you're going to need a lot of throughput.
|
||||
|
||||
Robert Caulk:
|
||||
The other is text embedding inference. This is a great server. A lot of vector databases will say, okay, we'll do all the embedding for you and we'll do all everything. But when you move to production scale, I'll talk a bit about this later. You need to be using micro service architecture, so it's not super smart to have your database bogged down with doing sorting out the embeddings and sorting out other things. So honestly, I'm a real big fan of single responsibility principle, and that's what Tei does for you. And it also does dynamic batching, which is great in this world where everything is heterogeneous lengths of what's coming in and what's going out. So it's great.
|
||||
@@ -129,7 +129,7 @@ Robert Caulk:
|
||||
The filters are huge. We're talking about real time filtering. We can't be searching on news articles from a month ago, two months ago, if the user is asking for a question that's related to the last 24 hours. So having that timestamp filtering and having it be efficient, which is what it is in Qdrant, is huge. Keyword filtering really opens up a massive realm of product opportunities for us. And then the sparse vectors, we hopped on this train immediately and are just seeing benefits. I don't want to say replacement of elasticsearch, but elasticsearch is using sparse vectors as well. So you can add splade into elasticsearch, and splade is great.
|
||||
|
||||
Robert Caulk:
|
||||
It's a really great alternative to BM 25. It's based on that Burt architecture, and that really opens up a lot of opportunities for filtering out keywords that are kind of useless to the search when the user uses the and a, and then there, these words that are less important splays a bit of a hybrid into semantics, but sparse retrieval. So it's really interesting. And then the idea of hybrid search with semantic and a sparse vector also opens up the ability to do ranking, and you got a higher quality product at the end, which is really the goal, right, especially in production. Point number four here, I would say, is probably one of the most important to us, because we're dealing in a world where latency is king, and being able to deploy Qdrant inside of the same cluster as all the other services. So we're just talking through the switch. That's huge. We're never getting bogged down by network.
|
||||
It's a really great alternative to BM 25. It's based on that architecture, and that really opens up a lot of opportunities for filtering out keywords that are kind of useless to the search when the user uses the and a, and then there, these words that are less important splays a bit of a hybrid into semantics, but sparse retrieval. So it's really interesting. And then the idea of hybrid search with semantic and a sparse vector also opens up the ability to do ranking, and you got a higher quality product at the end, which is really the goal, right, especially in production. Point number four here, I would say, is probably one of the most important to us, because we're dealing in a world where latency is king, and being able to deploy Qdrant inside of the same cluster as all the other services. So we're just talking through the switch. That's huge. We're never getting bogged down by network.
|
||||
|
||||
Robert Caulk:
|
||||
We're never worried about a cloud provider potentially getting overloaded or noisy neighbor problems, stuff like that, completely removed. And then you got high privacy, right. All the data is completely isolated from the external world. So this point number four, I'd say, is one of the biggest value adds for us. But then distributing deployment is huge because high availability is important, and deep storage, which when you're in the business of news archival, and that's one of our main missions here, is archiving the news forever. That's an ever growing database, and so you need a database that's going to be able to grow with you as your data grows. So what's the TLDR to this context? Engineering? Well, service orchestration is really just based on service orchestration in a very heterogeneous and parallel event driven environment. On the right side, we've got the user requests coming in.
|
||||
@@ -141,7 +141,7 @@ Robert Caulk:
|
||||
Open source projects like Qdrant, like Tei, like VLLM and Kubernetes, it's huge. Kubernetes is opening up doors for security and for latency. And of course, if you're going to be getting involved in this game, you got to find the strong DevOps. There's no escaping that. So let's step through kind of piece by piece and talk about flow Dapp. So that's our project. That's our open source project. We've spent about two years building this for our needs, and we're really excited because we did a public open sourcing maybe last week or the week before.
|
||||
|
||||
Robert Caulk:
|
||||
So finally, after all of our testing and rewrites and refactors, we're open. We're open for business. And it's running asknews app right now, and we're really excited for where it's going to go and how it's going to help other people orchestrate their clusters. Our goal and our priorities were highly paralyzed compute and we were running tests using all sorts of different executors, comparing them. So when you use Flowdapt, you can choose ray or dask. And that's key. Especially with vanilla Python, zero code changes, you don't need to know how ray or dask works. In the back end, floatapt is vanilla Python.
|
||||
So finally, after all of our testing and rewrites and refactors, we're open. We're open for business. And it's running asknews app right now, and we're really excited for where it's going to go and how it's going to help other people orchestrate their clusters. Our goal and our priorities were highly paralyzed compute and we were running tests using all sorts of different executors, comparing them. So when you use Flowdapt, you can choose ray or dask. And that's key. Especially with vanilla Python, zero code changes, you don't need to know how ray or dask works. In the back end, flowdapt is vanilla Python.
|
||||
|
||||
Robert Caulk:
|
||||
That was a key goal for us to ensure that we're optimizing how data is moving around the cluster. Automatic resource management this goes back to Ray and dask. They're helping manage the resources of the cluster, allocating a GPU to a task, or allocating multiple tasks to one GPU. These can come in very, very handy when you're dealing with very heterogeneous workloads like the ones that we discussed in those previous slides. For us, the biggest priority was ensuring rapid prototyping and debugging locally. When you're dealing with clusters of 1015 servers, 40 or 5100 with ray, honestly, ray just scales as far as you want. So when you're dealing with that big of a cluster, it's really imperative that what you see on your laptop is also what you are going to see once you deploy. And being able to debug anything you see in the cluster is big for us, we really found the need for easy cluster wide data sharing methods between tasks.
|
||||
@@ -150,19 +150,19 @@ Robert Caulk:
|
||||
So essentially what we've done is made it very easy to get and put values. And so this makes it extremely easy to move data and share data between tasks and make it highly available and stay in cluster memory or persist it to disk, so that when you do the inevitable version update or debug, you're reloading from a persisted state in the real time. News business scheduling is huge. Scheduling, making sure that various workflows are scheduled at different points and different periods or frequencies rather, and that they're being scheduled correctly, and that their triggers are triggering exactly what you need when you need it. Huge for real time. And then one of our biggest selling points, if you will, for this project is Kubernetes style. Everything. Our goal is everything's Kubernetes style, so that if you're coming from Kubernetes, everything's familiar, everything's resource oriented.
|
||||
|
||||
Robert Caulk:
|
||||
We even have our own flow ectyl, which would be the Kubectl style command schemas. A lot of what we've done is ensuring deployment cycle efficiency here. So the goal is that flowdapt can schedule everything and manage all these services for you, create workflows. But why these services? For this particular use case, I'll kind of skip through quickly. I know I'm kind of running out of time here, but of course you're going to need some proprietary remote models. That's just how it works. You're going to of course share that load with on premise llms to reduce cost and to have some reasoning engine on premise. But there's obviously advantages and disadvantages to these.
|
||||
We even have our own flowectl, which would be the Kubectl style command schemas. A lot of what we've done is ensuring deployment cycle efficiency here. So the goal is that flowdapt can schedule everything and manage all these services for you, create workflows. But why these services? For this particular use case, I'll kind of skip through quickly. I know I'm kind of running out of time here, but of course you're going to need some proprietary remote models. That's just how it works. You're going to of course share that load with on premise llms to reduce cost and to have some reasoning engine on premise. But there's obviously advantages and disadvantages to these.
|
||||
|
||||
Robert Caulk:
|
||||
I'm not going to go through them. I'm happy to make these slides available, and you're welcome to kind of parse through the details. Yeah, for sure. You need to start thinking about persistence and search and making sure those services are robust. That's where Qdrant comes into play. And we found that the all in one solutions kind of sacrifice performance for convenience, or sacrifice accuracy for convenience, but it really wasn't for us. We'd rather just orchestrate it ourselves and let Qdrant do what Qdrant does, instead of kind of just hope that an all in one solution is handling it for us and that allows for modularity performance. And we'll dump Qdrant if we want to.
|
||||
|
||||
Robert Caulk:
|
||||
Probably we won't. Or we'll dump minio if we need to, or we'll swap out for whatever replaces bllm. Trying to keep things modular so that future engineers are able to adapt with the tech that's just blowing up and exploding right now. Right. The last thing to talk about here in a production scale environment is really minimizing the latency. I touched on this with Kubernetes ensuring that these services are sitting on the same network, and that is huge. But that talks about decommunication latency. But when you start talking about getting hit with a ton of traffic, production scale, tons of people asking a question all simultaneously, and you needing to go hit a variety of services, well, this is where you really need to isolate that to an asynchronous environment.
|
||||
Probably we won't. Or we'll dump if we need to, or we'll swap out for whatever replaces vllm. Trying to keep things modular so that future engineers are able to adapt with the tech that's just blowing up and exploding right now. Right. The last thing to talk about here in a production scale environment is really minimizing the latency. I touched on this with Kubernetes ensuring that these services are sitting on the same network, and that is huge. But that talks about decommunication latency. But when you start talking about getting hit with a ton of traffic, production scale, tons of people asking a question all simultaneously, and you needing to go hit a variety of services, well, this is where you really need to isolate that to an asynchronous environment.
|
||||
|
||||
Robert Caulk:
|
||||
And of course, if you could write this all in Golang, that's probably going to be your best bet for us. We have some services written in Golang, but predominantly, especially the endpoints that the ML engineers need to work with. We're using fast API on pydantic and honestly, it's powerful. Pydantic V 2.0 now runs on Rust, and as anyone in the Qdrant community knows, Rust is really valuable when you're dealing with highly parallelized environments that require high security and protections for immutability and atomicity. Forgive me for the pronunciation, that kind of sums up the production scale talk, and I'm happy to answer questions. I love diving into this sort of stuff. I do have some just general thoughts on why startups are so much more well positioned right now than some of these incumbents, and I'll just do kind of a quick run through, less than a minute just to kind of get it out there. We can talk about it, see if we agree or disagree.
|
||||
|
||||
Robert Caulk:
|
||||
But you touched on it, Demetrius, in the introduction, which was the best practices have not been established. That's it. That is why startups have such a big advantage. And the reason they're not established is because, well, the new paradigm of technology is just underexplored. We don't really know what the limits are and how to properly handle these things. And that's huge. Meanwhile, some of these incumbents, they're dealing with all sorts of limitations and resistance to change and stuff, and then just market expectations for incumbents maintaining these kind of legacy products and trying to keep them hobbling along on this old tech. In my opinion, startups, you got your reasoning engine building everything around a reasoning engine, using that reasoning engine for every aspect of your system to really open up the adaptivity of your product.
|
||||
But you touched on it, Demetrios, in the introduction, which was the best practices have not been established. That's it. That is why startups have such a big advantage. And the reason they're not established is because, well, the new paradigm of technology is just underexplored. We don't really know what the limits are and how to properly handle these things. And that's huge. Meanwhile, some of these incumbents, they're dealing with all sorts of limitations and resistance to change and stuff, and then just market expectations for incumbents maintaining these kind of legacy products and trying to keep them hobbling along on this old tech. In my opinion, startups, you got your reasoning engine building everything around a reasoning engine, using that reasoning engine for every aspect of your system to really open up the adaptivity of your product.
|
||||
|
||||
Robert Caulk:
|
||||
And okay, I won't put elasticsearch in the incumbent world. I'll keep elasticsearch in the middle. I understand it still has a lot of value, but some of these vendor lock ins, not a huge fan of. But anyway, that's it. That's kind of all I have to say. But I'm happy to take questions or chat a bit.
|
||||
@@ -189,7 +189,7 @@ Robert Caulk:
|
||||
This is another logistical point that we think needs to get sorted properly and there's a few layers to it. So for us, as we're parsing that data coming in from Newscatcher, so newscatcher is doing a good job of always feeding the latest buckets to us. Sometimes one will be kind of arrive, but generally speaking, it's always the latest news. So we're taking five minute buckets, and then with those buckets, we're going through and doing all of our enrichment on that, adding it to Qdrant. And that is the point where we use that timestamp filtering, which is such an important point. So in the metadata of Qdrant, we're using the range filter, which is where we call that the timestamp filter, but it's really range filter, and that helps. So when we're going back to update things, we're sorting and ensuring that we're filtering out only what we haven't seen.
|
||||
|
||||
Demetrios:
|
||||
Okay, that makes complete sense. And basically you could generalize this to something like what I was talking to with people yesterday about, which was, hey, I've got an HR policy that gets updated every other month or every quarter, and I want to make sure that if my HR chat bot is telling people what their vacation policy is, it's pulling from the most recent HR policy. So how do I make sure and do that? And how do I make sure that my vector database isn't like a landmine where it's pulling any information, but we don't necessarily have that control to be able to pull the correct information? And this comes down to that retrieval evaluation, which is such a hot topic, too.
|
||||
Okay, that makes complete sense. And basically you could generalize this to something like what I was talking to with people yesterday about, which was, hey, I've got an HR policy that gets updated every other month or every quarter, and I want to make sure that if my HR chatbot is telling people what their vacation policy is, it's pulling from the most recent HR policy. So how do I make sure and do that? And how do I make sure that my vector database isn't like a landmine where it's pulling any information, but we don't necessarily have that control to be able to pull the correct information? And this comes down to that retrieval evaluation, which is such a hot topic, too.
|
||||
|
||||
Robert Caulk:
|
||||
That's true. No, I think that's a key piece of the puzzle. Now, in that particular example, maybe you actually want to go in and start cleansing a bit, your database, just to make sure if it's really something you're never going to need again. You got to get rid of it. This is a piece I didn't add to the presentation, but it's tangential. You got to keep multiple databases and you got to making sure to isolate resources and cleaning out a database, especially in real time. So ensuring that your database is representative of what you want to be searching on. And you can do this with collections too, if you want.
|
||||
@@ -198,10 +198,10 @@ Robert Caulk:
|
||||
But we find there's sometimes a good opportunity to isolate resources in that sense, 100%.
|
||||
|
||||
Demetrios:
|
||||
So, another question that I had for you was, I noticed Mongo was in the stack. Why did you not just use the Mongo vector option? Is it because of what you were mentioning, where it's like, yeah, you have these all in one options, but you sacrifice that performance for the convenience?
|
||||
So, another question that I had for you was, I noticed Mongo was in the stack. Why did you not just use the Mongo vector option? Is it because of what you were mentioning, where it's like, yeah, you have these all-in-one options, but you sacrifice that performance for the convenience?
|
||||
|
||||
Robert Caulk:
|
||||
We didn't test that, to be honest, I can't say. All I know is we tested weavyt, we tested one other, and I just really like. Although I was going to say I like that it's written in rust, although I believe Mongo is also written in rust, if I'm not mistaken. But for us, the document DB is more of a representation of state and what's happening, especially for our configurations and workflows. Meanwhile, we really like keeping and relying on Qdrant and all the features. Qdrant is updating, so, yeah, I'd say single responsibility principle is key to that. But I saw some chat in Qdrant discord about this, which I think the only way to use vector is actually to use their cloud offering, if I'm not mistaken. Do you know about this?
|
||||
We didn't test that, to be honest, I can't say. All I know is we tested weaviate, we tested one other, and I just really like. Although I was going to say I like that it's written in rust, although I believe Mongo is also written in rust, if I'm not mistaken. But for us, the document DB is more of a representation of state and what's happening, especially for our configurations and workflows. Meanwhile, we really like keeping and relying on Qdrant and all the features. Qdrant is updating, so, yeah, I'd say single responsibility principle is key to that. But I saw some chat in Qdrant discord about this, which I think the only way to use vector is actually to use their cloud offering, if I'm not mistaken. Do you know about this?
|
||||
|
||||
Demetrios:
|
||||
Yeah, I think so, too.
|
||||
|
||||
@@ -1,16 +1,16 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: Talk with YouTube without paying a cent - Francesco Saverio Zuppichini |
|
||||
Vector Space Talks
|
||||
slug: youtube-without-paying-cent
|
||||
short_description: Dive deep into the tech world as Francesco shares his
|
||||
insights and processes on coding innovative solutions.
|
||||
description: Francesco Zuppichini intricately outlines the process of converting
|
||||
YouTube video subtitles into searchable vector databases, leveraging tools
|
||||
like YouTube DL and Hugging Face, and addressing the challenges of coding
|
||||
without conventional frameworks in machine learning engineering.
|
||||
short_description: A sneak peek into the tech world as Francesco shares his
|
||||
ideas and processes on coding innovative solutions.
|
||||
description: Francesco Zuppichini outlines the process of converting YouTube
|
||||
video subtitles into searchable vector databases, leveraging tools like
|
||||
YouTube DL and Hugging Face, and addressing the challenges of coding without
|
||||
conventional frameworks in machine learning engineering.
|
||||
preview_image: /blog/from_cms/francesco-saverio-zuppichini-bp-cropped.png
|
||||
date: 2024-03-08T10:12:59.752Z
|
||||
date: 2024-03-27T12:37:55.643Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -35,11 +35,11 @@ Francesco Saverio Zuppichini is a Senior Full Stack Machine Learning Engineer at
|
||||
|
||||
Curious about transforming YouTube content into searchable elements? Francesco Zuppichini unpacks the journey of coding a RAG by using subtitles as input, harnessing technologies like YouTube DL, Hugging Face, and Qdrant, while debating framework reliance and the fine art of selecting the right software tools.
|
||||
|
||||
Here are some insights from this great episode:
|
||||
Here are some insights from this episode:
|
||||
|
||||
1. **Behind the Code**: Francesco unravels how to create a RAG using YouTube videos. Get ready to geek out on the nuts and bolts that make this magic happen.
|
||||
2. **Vector Voodoo**: Ever wonder how embedding vectors carry out their similarity searches? Francesco's got you covered with his brilliant explanation of vector databases and the mind-bending distance method that seeks out those matches.
|
||||
3. **Function over Class**: The debate is as old as stardust. Francesco shares why he prefers using functions over classes for better code organization and demonstrates how this approach crystallizes when running language models with Ollama.
|
||||
3. **Function over Class**: The debate is as old as stardust. Francesco shares why he prefers using functions over classes for better code organization and demonstrates how this approach solidifies when running language models with Ollama.
|
||||
4. **Metadata Magic**: Find out how metadata isn't just a sidekick but plays a pivotal role in the realm of Qdrant and RAGs. Learn why Francesco values metadata as payload and the challenges it presents in developing domain-specific applications.
|
||||
5. **Tool Selection Tips**: Deciding on the right software tool can feel like navigating an asteroid belt. Francesco shares his criteria—ease of installation, robust documentation, and a little help from friends—to ensure a safe landing.
|
||||
|
||||
@@ -132,10 +132,10 @@ Demetrios:
|
||||
Yes, for sure.
|
||||
|
||||
Francesco Zuppichini:
|
||||
That's perfect. Okay, so today we're going to talk about talk with YouTube without paying a cent, no framework bs. So the goal of today is to showcase how to code a RAG given as an input a YouTube video without using any framework like language, et cetera, et cetera. And I want to show you that it's straightforward, using a bunch of technologies and Qdrants as well. And you can do all of this without actually pay to any service. Right. So we are going to run our PetrodB locally and also the language model. We are going to run our machines.
|
||||
That's perfect. Okay, so today we're going to talk about talk with YouTube without paying a cent, no framework bs. So the goal of today is to showcase how to code a RAG given as an input a YouTube video without using any framework like language, et cetera, et cetera. And I want to show you that it's straightforward, using a bunch of technologies and Qdrants as well. And you can do all of this without actually pay to any service. Right. So we are going to run our PEDro DB locally and also the language model. We are going to run our machines.
|
||||
|
||||
Francesco Zuppichini:
|
||||
And yeah, it's going to be a technical talk, so I will kind of guide you through the code. Feel free to interrupt me at any time if you have questions, if you want to ask why I did that, et cetera, et cetera. So very quickly, before we get started, I just want you not to introduce myself. So yeah, senior full stack machine engineer. That's just a bunch of funny work to basically say that I do a little bit of everything. Start. So when I was working, I start as computer vision engineer, I work at PwC, then a bunch of startups, and now I sold my soul to insurance companies working at insurance. And before I was doing computer vision, now I'm doing due to chat CPT, hyper language model, I'm doing more of that.
|
||||
And yeah, it's going to be a technical talk, so I will kind of guide you through the code. Feel free to interrupt me at any time if you have questions, if you want to ask why I did that, et cetera, et cetera. So very quickly, before we get started, I just want you not to introduce myself. So yeah, senior full stack machine engineer. That's just a bunch of funny work to basically say that I do a little bit of everything. Start. So when I was working, I start as computer vision engineer, I work at PwC, then a bunch of startups, and now I sold my soul to insurance companies working at insurance. And before I was doing computer vision, now I'm doing due to Chat GPT, hyper language model, I'm doing more of that.
|
||||
|
||||
Francesco Zuppichini:
|
||||
But I'm always involved in bringing the full product together. So from zero to something that is deployed and running. So I always be interested in web dev. I can also do website servers, a little bit of infrastructure as well. So now I'm just doing a little bit of everything. So this is why there is full stack there. Yeah. Okay, let's get started to something a little bit more interesting than myself.
|
||||
@@ -165,13 +165,13 @@ Francesco Zuppichini:
|
||||
Wonderful. Okay, so in order to get the embedding. So to translate from text to vectors, right, so we're going to use hugging face just an embedding model so we can actually get some vectors. Then as soon as we got our vectors, we need to store and search them. So we're going to use our beloved Qdrant to do so. We also need to keep a little bit of stage right because we need to know which video we have processed so we don't redo the old embeddings and the storing every time we see the same video. So for this part, I'm just going to use SQLite, which is just basically an SQL database in just a file. So very easy to use, very kind of lightweight, and it's only your computer, so it's safe to run the language model.
|
||||
|
||||
Francesco Zuppichini:
|
||||
We're going to use Olama. That is a very simple way and very well done way to just get a language model that is running on your computer. And you can also call it using the OpenAI Python library because they have implemented the same endpoint as. It's like, it's super convenient, super easy to use. If you already have some code that is calling OpenAI, you can just run a different language model using Olama. And you just need to basically change two lines of code. So what we're going to do, basically, I'm going to take a video. So here it's a video from Fireship IO.
|
||||
We're going to use Ollama. That is a very simple way and very well done way to just get a language model that is running on your computer. And you can also call it using the OpenAI Python library because they have implemented the same endpoint as. It's like, it's super convenient, super easy to use. If you already have some code that is calling OpenAI, you can just run a different language model using Ollama. And you just need to basically change two lines of code. So what we're going to do, basically, I'm going to take a video. So here it's a video from Fireship IO.
|
||||
|
||||
Francesco Zuppichini:
|
||||
We're going to run our command line and we're going to ask some questions. Now, if you can still, in theory, you should be able to see my full screen. Yeah. So very quickly to showcase that to you, I already processed this video from the good sound YouTube channel and I have already here my command line. So I can already kind of see, you know, I can ask a question like what is the contact size of Germany? And we're going to get the reply. Yeah. And here we're going to get a reply. And now I want to walk you through how you can do something similar.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Now, the goal is not to create the best rack in the world. It's just to showcase like show zero to something that is actually working. How you can do that in a fully local way without using any framework so you can really understand what's going on under the hood. Because I think a lot of people, they try to copy, to just copy and paste stuff on langchain and then they end up in a situation when they need to change something, but they don't really know where the stuff is. So this is why I just want to just show like Windfield zero to hero. So the first step will be I get a YouTube video and now I need to get the subtitle. So you could actually use a model to take the audio from the video and get the text. Like a whisper model from OpenAI, for example.
|
||||
Now, the goal is not to create the best rack in the world. It's just to showcase like show zero to something that is actually working. How you can do that in a fully local way without using any framework so you can really understand what's going on under the hood. Because I think a lot of people, they try to copy, to just copy and paste stuff on Langchain and then they end up in a situation when they need to change something, but they don't really know where the stuff is. So this is why I just want to just show like Windfield zero to hero. So the first step will be I get a YouTube video and now I need to get the subtitle. So you could actually use a model to take the audio from the video and get the text. Like a whisper model from OpenAI, for example.
|
||||
|
||||
Francesco Zuppichini:
|
||||
In this case, we are taking advantage that YouTube allow people to upload subtitles and YouTube will automatically generate the subtitles. So here using YouTube dial, I'm just going to get my video URL. I'm going to set up a bunch of options like the format they want, et cetera, et cetera. And then basically I'm going to download and get the subtitles. And they look something like this. Let me show you an example. Something similar to this one, right? We have the timestamps and we do have all text inside. Now the next step.
|
||||
@@ -195,7 +195,7 @@ Francesco Zuppichini:
|
||||
You just need to do this once. I was very lazy so I just assumed that if this is going to fail, it means that it's because I've already created a collection. So I'm just going to pass it and call it a day. Okay, so this is basically all the preprocess this setup you need to do to have your Qdrant ready to store and search vectors. To store vectors. Straightforward, very straightforward as well. Just need again the client. So the connection to the database here I'm passing my embedding so sentence transformer model and I'm passing my chunks as a list of documents.
|
||||
|
||||
Francesco Zuppichini:
|
||||
So documents in my code is just a type dict that will contain just this metadata here. Very simple. It's similar to Lang chain here. I just have attacked it because it's lightweight. To store them we call the upload records function. We encode them here. There is a little bit of bad variable names from my side which I replacing that. So you shouldn't do that.
|
||||
So documents in my code is just a type that will contain just this metadata here. Very simple. It's similar to Lang chain here. I just have attacked it because it's lightweight. To store them we call the upload records function. We encode them here. There is a little bit of bad variable names from my side which I replacing that. So you shouldn't do that.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Apologize about that and you just send the records. Another very cool thing about Qdrant. So the second things that I really like is that they have types for what you send through the library. So this models record is a Qdrant type. So you use it and you know immediately. So what you need to put inside. So let me give you an example. Right? So assuming that I'm programming, right, I'm going to say model record bank.
|
||||
@@ -213,13 +213,13 @@ Francesco Zuppichini:
|
||||
We need to recreate to embed in a vector and then we need to compare with the vectors in the vector Db using a distance method, in this case considered similarity in order to get the right matches right, the closest one in our vector DB, in our vector search base. So passing a query string, I'm passing a video id and I pass in a label. So how many hits I want to get from the metadb. Now to create a filter again you're going to use the model package from the Qdrant framework. So here I'm just creating a filter class for the model and I'm saying okay, this filter must match this key, right? So metadata video id with this video id. So when we search, before we do the similarity search, we are going to filter away all the vectors that are not from that video. Wonderful. Now super easy as well.
|
||||
|
||||
Francesco Zuppichini:
|
||||
We just call the DB search, right pass. Our collection name here is star coded. Apologies about that, I think I forgot to put the right global variable our coded, we create a query, we set the limit, we pass the query filter, we get the it back as a dictionary in the payload field of each it and we recreate our document a dictionary. I have types, right? So I know what this function is going to return. Now if you were to use a framework, right this part, it will be basically the same thing. If I were to use lamb chain and I want to specify a filter, I would have to write the same amount of code. So most of the times you don't really need to use a framework. One thing that is nice about not using a framework here is that I add control on the indexes.
|
||||
We just call the DB search, right pass. Our collection name here is star coded. Apologies about that, I think I forgot to put the right global variable our coded, we create a query, we set the limit, we pass the query filter, we get the it back as a dictionary in the payload field of each it and we recreate our document a dictionary. I have types, right? So I know what this function is going to return. Now if you were to use a framework, right this part, it will be basically the same thing. If I were to use langchain and I want to specify a filter, I would have to write the same amount of code. So most of the times you don't really need to use a framework. One thing that is nice about not using a framework here is that I add control on the indexes.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Long chain, for instance, will create the indexes only while you call a classmate like from document. And that is kind of cumbersome because sometimes I wasn't quoting bugs in which I was not understanding why one index was created before, after, et cetera, et cetera. So yes, just try to keep things simple and not always write on frameworks. Wonderful. Now I have a way to ask a query to get back the relative parts from that video. Now we need to translate this list of chunks to something that we can read as human. Before we do that, I was almost going to forget we need to keep state. Now, one of the last missing part is something in which I can store data.
|
||||
Lang chain, for instance, will create the indexes only while you call a classmate like from document. And that is kind of cumbersome because sometimes I wasn't quoting bugs in which I was not understanding why one index was created before, after, et cetera, et cetera. So yes, just try to keep things simple and not always write on frameworks. Wonderful. Now I have a way to ask a query to get back the relative parts from that video. Now we need to translate this list of chunks to something that we can read as human. Before we do that, I was almost going to forget we need to keep state. Now, one of the last missing part is something in which I can store data.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Here I just have a setup function in which I'm going to create an SQL lite database, create a table called videos in which I have an id and a title. So later I can check, hey, is this video already in my database? Yes. I don't need to process that. I can just start immediately to q a on that video. If not, I'm going to do the chunking and embeddings. Got a couple of functions here to get video from Db to save video from and to save video to Db. So notice now I only use functions. I'm not using classes here.
|
||||
Here I just have a setup function in which I'm going to create an SQL lite database, create a table called videos in which I have an id and a title. So later I can check, hey, is this video already in my database? Yes. I don't need to process that. I can just start immediately to QA on that video. If not, I'm going to do the chunking and embeddings. Got a couple of functions here to get video from Db to save video from and to save video to Db. So notice now I only use functions. I'm not using classes here.
|
||||
|
||||
Francesco Zuppichini:
|
||||
I'm not a fan of object writing programming because it's very easy to kind of reach inheritance health in which we have like ten levels of inheritance. And here if a function needs to have state, here we do need to have state because we need a connection. So I will just have a function that initialize that state. I return tat to me, and me as a caller, I'm just going to call it and pass my state. Very simple tips allow you really to divide your code properly. You don't need to think about is my class to couple with another class, et cetera, et cetera. Very simple, very effective. So what I suggest when you're coding, just start with function and share states across just pass down state.
|
||||
@@ -228,13 +228,13 @@ Francesco Zuppichini:
|
||||
And when you realize that you can cluster a lot of function together with a common behavior, you can go ahead and put state in a class and have key function as methods. So try to not start first by trying to understand which class I need to use around how I connect them, because in my opinion it's just a waste of time. So just start with function and then try to cluster them together if you need to. Okay, last part, the juicy part as well. Language models. So we need the language model. Why do we need the language model? Because I'm going to ask a question, right. I'm going to get a bunch of relevant chunks from a video and the language model.
|
||||
|
||||
Francesco Zuppichini:
|
||||
It needs to answer that to me. So it needs to get information from the chunks and reply that to me using that information as a context. To run language model, the easiest way in my opinion is using Olama. There are a lot of models that are available. I put a link here and you can also bring your own model. There are a lot of videos and tutorial how to do that. You run this command as soon as you install it on Linux. It's a one line to install o llama.
|
||||
It needs to answer that to me. So it needs to get information from the chunks and reply that to me using that information as a context. To run language model, the easiest way in my opinion is using Ollama. There are a lot of models that are available. I put a link here and you can also bring your own model. There are a lot of videos and tutorial how to do that. You run this command as soon as you install it on Linux. It's a one line to install Ollama.
|
||||
|
||||
Francesco Zuppichini:
|
||||
You run this command here, it's going to download Mistral seven B very good model and run it on your gpu if you have one, or your cpu if you don't have a gpu, run it on GPU. Here you can see it yet. It's around 6gb. So even with a low tier gpu, you should be able to run a seven minute model on your gpu. Okay, so this is the prompt just for also to show you how easy is this, this prompt was just very lazy. Copy and paste from langchain source code here prompt use the following piece of context to answer the question at the end. Blah blah blah variable to inject the context inside question variable to get question and then we're going to get an answer. How do we call it? Is it peasy? I have a function here called getanswer passing a bunch of stuff, passing also the OpenAI from the OpenAI Python package model client passing a question, passing a vdb, my Db client, my embeddings, reading my prompt, getting my matching documents, calling the search function we have just seen before, creating my context.
|
||||
You run this command here, it's going to download Mistral 7B very good model and run it on your gpu if you have one, or your cpu if you don't have a gpu, run it on GPU. Here you can see it yet. It's around 6gb. So even with a low tier gpu, you should be able to run a seven minute model on your gpu. Okay, so this is the prompt just for also to show you how easy is this, this prompt was just very lazy. Copy and paste from langchain source code here prompt use the following piece of context to answer the question at the end. Blah blah blah variable to inject the context inside question variable to get question and then we're going to get an answer. How do we call it? Is it easy? I have a function here called getanswer passing a bunch of stuff, passing also the OpenAI from the OpenAI Python package model client passing a question, passing a vdb, my DB client, my embeddings, reading my prompt, getting my matching documents, calling the search function we have just seen before, creating my context.
|
||||
|
||||
Francesco Zuppichini:
|
||||
So just joining the text in the chunks on a new line, calling the format function in Python. As simple as that. Just calling the format function in Python because the format function will look at a string and kitty will inject variables that match inside these parentheses. Passing context passing question using the Openi model client APIs and getting a reply back. Super easy. And here I'm returning the reply from the language model and also the list of documents. So this should be documents. I think I did a mistake.
|
||||
So just joining the text in the chunks on a new line, calling the format function in Python. As simple as that. Just calling the format function in Python because the format function will look at a string and kitty will inject variables that match inside these parentheses. Passing context passing question using the OpenAI model client APIs and getting a reply back. Super easy. And here I'm returning the reply from the language model and also the list of documents. So this should be documents. I think I did a mistake.
|
||||
|
||||
Francesco Zuppichini:
|
||||
When I copy and paste this to get this image and we are done right. We have a way to get some answers from a video by putting everything together. This can seem scary because there is no comment here, but I can show you tson code. I think it's easier so I can highlight stuff. I'm creating my embeddings, I'm getting my database, I'm getting my vector DB login, some stuff I'm getting my model client, I'm getting my vid. So here I'm defining the state that I need. You don't need comments because I get it straightforward. Like here I'm getting the vector db, good function name.
|
||||
@@ -288,7 +288,7 @@ Francesco Zuppichini:
|
||||
Nice to everyone by the way.
|
||||
|
||||
Demetrios:
|
||||
So from my side, I'm wondering, do you have any specific design decisions criteria that you use when you are building out your stack? Like you chose Mistral, you chose Olama, you chose Qdrant. It sounds like with Qdrant you did some testing and you appreciated the capabilities. With Qdrant, was it similar with Olama and Mistral?
|
||||
So from my side, I'm wondering, do you have any specific design decisions criteria that you use when you are building out your stack? Like you chose Mistral, you chose Ollama, you chose Qdrant. It sounds like with Qdrant you did some testing and you appreciated the capabilities. With Qdrant, was it similar with Ollama and Mistral?
|
||||
|
||||
Francesco Zuppichini:
|
||||
So my test is how long it's going to take to install that tool. If it's taking too much time and it's hard to install because documentation is bad, so that it's a red flag, right? Because if it's hard to install and documentation is bad for the installation, that's the first thing people are going to read. So probably it's not going to be great for something down the road to use Olama. It took me two minutes, took me two minutes, it was incredible. But just install it, run it and it was done. Same thing with Qualent as well and same thing with the hacking phase library. So to me, usually as soon as if I see that something is easy to install, that's usually means that is good. And if the documentation to install it, it's good.
|
||||
@@ -318,7 +318,7 @@ Sabrina Aquino:
|
||||
Yeah, that's a great explanation of collections. And I do love your approach of having everything locally and having everything in a structured way that you can really understand what you're doing. And I know you mentioned sometimes frameworks are not necessary. And I wonder also from your side, when do you think a framework would be necessary and does it have to do with scaling? What do you think?
|
||||
|
||||
Francesco Zuppichini:
|
||||
So that's a great question. So what frameworks in theory should give you is good interfaces, right? So a good interface means that if I'm following that interface, I know that I can always call something that implements that interface in the same way. Like for instance in Lanchain, if I call a betterdb, I can just swap the betterdb and I can call it in the same way. If the interfaces are good, the framework is useful. If you know that you are going to change stuff. In my case, I know from the beginning that I'm going to use Qdrant, I'm going to use Holama, and I'm going to use SQL lite. So why should I go to the hello reading framework documentation? I install libraries, and then you need to install a bunch of packages from the framework that you don't even know why you need them. Maybe you have a conflict package, et cetera, et cetera.
|
||||
So that's a great question. So what frameworks in theory should give you is good interfaces, right? So a good interface means that if I'm following that interface, I know that I can always call something that implements that interface in the same way. Like for instance in Langchain, if I call a betterdb, I can just swap the betterdb and I can call it in the same way. If the interfaces are good, the framework is useful. If you know that you are going to change stuff. In my case, I know from the beginning that I'm going to use Qdrant, I'm going to use Ollama, and I'm going to use SQL lite. So why should I go to the hello reading framework documentation? I install libraries, and then you need to install a bunch of packages from the framework that you don't even know why you need them. Maybe you have a conflict package, et cetera, et cetera.
|
||||
|
||||
Francesco Zuppichini:
|
||||
If you know ready. So what you want to do then just code it and call it a day? Like in this case, I know I'm not going to change the vector DB. If you think that you're going to change something, even if it's a simple approach, it's fair enough, simple to change stuff. Like I will say that if you know that you want to change your vector DB providers, either you define your own interface or you use a framework with an already defined interface. But be careful because right too much on framework will. First of all, basically you don't know what's going on inside the hood for launching because it's so kudos to them. They were the first one. They are very smart people, et cetera, et cetera.
|
||||
|
||||
@@ -0,0 +1,263 @@
|
||||
---
|
||||
draft: true
|
||||
title: "Teaching Vector Databases at Scale - Alfredo Deza | Vector Space Talks #019"
|
||||
slug: teaching-vector-db-at-scale
|
||||
short_description: Alfredo Deza tackles AI teaching, the intersection of
|
||||
technology and academia, and the value of consistent learning.
|
||||
description: Alfredo Deza discusses the practicality of machine learning
|
||||
operations, highlighting how personal interest in topics like wine datasets
|
||||
enhances engagement, while reflecting on the synergies between his athletic
|
||||
discipline and the persistent, straightforward approach required for
|
||||
effectively educating on vector databases and large language models.
|
||||
preview_image: /blog/from_cms/alfredo-deza-bp-cropped.png
|
||||
date: 2024-04-02T22:57
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
- Vector Search
|
||||
- Retrieval Augmented Generation
|
||||
- Vector Space Talks
|
||||
- Coursera
|
||||
---
|
||||
> *"So usually I get asked, why are you using Qdrant? What's the big deal? Why are you picking these over all of the other ones? And to me it boils down to, aside from being renowned or recognized, that it works fairly well. There's one core component that is critical here, and that is it has to be very straightforward, very easy to set up so that I can teach it, because if it's easy, well, sort of like easy to or straightforward to teach, then you can take the next step and you can make it a little more complex, put other things around it, and that creates a great development experience and a learning experience as well.”*\
|
||||
— Alfredo Deza
|
||||
>
|
||||
|
||||
Alfredo is a software engineer, speaker, author, and former Olympic athlete working in Developer Relations at Microsoft. He has written several books about programming languages and artificial intelligence and has created online courses about the cloud and machine learning.
|
||||
|
||||
He currently is an Adjunct Professor at Duke University, and as part of his role, works closely with universities around the world like Georgia Tech, Duke University, Carnegie Mellon, and Oxford University where he often gives guest lectures about technology.
|
||||
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/4HFSrTJWxl7IgQj8j6kwXN?si=99H-p0fKQ0WuVEBJI9ugUw), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/3l6F6A_It0Q?feature=shared).***
|
||||
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/3l6F6A_It0Q?si=cFZGAh7995iHilcY" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
<iframe src="https://podcasters.spotify.com/pod/show/qdrant-vector-space-talk/embed/episodes/Teaching-Vector-Databases-at-Scale---Alfredo-Deza--Vector-Space-Talks-019-e2hhjlo/a-ab3qp7u" height="102px" width="400px" frameborder="0" scrolling="no"></iframe>
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
How does a former athlete such as Alfredo Deza end up in this AI and Machine Learning industry? That’s what we’ll find out in this episode of Vector Space Talks. Let’s understand how his background as an olympian offers a unique perspective on consistency and discipline that's a real game-changer in this industry.
|
||||
|
||||
Here are some things you’ll discover from this episode:
|
||||
|
||||
1. **The Intersection of Teaching and Tech:** Alfredo discusses on how to effectively bridge the gap between technical concepts and student understanding, especially when dealing with complex topics like vector databases.
|
||||
2. **Simplified Learning:** Dive into Alfredo's advocacy for simplicity in teaching methods, mirroring his approach with Qdrant and the potential for a Rust in-memory implementation aimed at enhancing learning experiences.
|
||||
3. **Beyond the Titanic Dataset:** Discover why Alfredo prefers to teach with a wine dataset he developed himself, underscoring the importance of using engaging subject matter in education.
|
||||
4. **AI Learning Acceleration:** Alfredo discusses the struggle universities face to keep pace with AI advancements and how online platforms can offer a more up-to-date curriculum.
|
||||
5. **Consistency is Key:** Alfredo draws parallels between the discipline required in high-level athletics and the ongoing learning journey in AI, zeroing in on his mantra, “There is no secret” to staying consistent.
|
||||
|
||||
> Fun Fact: Alfredo tells the story of athlete Dick Fosbury's invention of the Fosbury Flop to highlight the significance of teaching simplicity.
|
||||
>
|
||||
|
||||
## Show notes:
|
||||
|
||||
00:00 Teaching machine learning, Python to graduate students.\
|
||||
06:03 Azure AI search service simplifies teaching, Qdrant facilitates learning.\
|
||||
10:49 Controversy over high jump style.\
|
||||
13:18 Embracing past for inspiration, emphasizing consistency.\
|
||||
15:43 Consistent learning and practice lead to success.\
|
||||
20:26 Teaching SQL uses SQLite, Rust has limitations.\
|
||||
25:21 Online platforms improve and speed up education.\
|
||||
29:24 Duke and Coursera offer specialized language courses.\
|
||||
31:21 Passion for wines, creating diverse dataset.\
|
||||
35:00 Encouragement for vector db discussion, wrap up.
|
||||
|
||||
## More Quotes from Alfredo:
|
||||
|
||||
*"Qdrant makes it straightforward. We use it in-memory for my classes and I would love to see something similar setup in Rust to make teaching even easier.”*\
|
||||
— Alfredo Deza
|
||||
|
||||
*"Retrieval augmented generation is kind of like having an open book test. So the large language model is the student, and they have an open book so they can see the answers and then repackage that into their own words and provide an answer.”*\
|
||||
— Alfredo Deza
|
||||
|
||||
*"With Qdrant, I appreciate that the use of the Python API is so simple. It avoids the complexity that comes from having a back-end system like in Rust where you need an actual instance of the database running.”*\
|
||||
— Alfredo Deza
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
What is happening? Everyone, welcome back to another vector space talks. I am Demetrios, and I am joined today by good old Sabrina. Where you at, Sabrina? Hello?
|
||||
|
||||
Sabrina Aquino:
|
||||
Hello, Demetrios. I'm from Brazil. I'm in Brazil right now. I know that you are traveling currently.
|
||||
|
||||
Demetrios:
|
||||
Where are you? At Kubecon in Paris. And it has been magnificent. But I could not wait to join the session today because we've got Alfredo coming at us.
|
||||
|
||||
Alfredo Deza:
|
||||
What's up, dude? Hi. How are you?
|
||||
|
||||
Demetrios:
|
||||
I'm good, man. It's been a while. I think the last time that we chatted was two years ago, maybe right before your book came out. When did the book come out?
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, something like that. I would say a couple of years ago. Yeah. I wrote, co authored practical machine learning operations with no gift. And it was published on O'Reilly.
|
||||
|
||||
Demetrios:
|
||||
Yeah. And that was, I think, two years ago. So you've been doing a lot of stuff since then. Let's be honest, you are maybe one of the most active men on the Internet. I always love seeing what you're doing. You're bringing immense value to everything that you touch. I'm really excited to be able to chat with you for this next 30 minutes.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, of course.
|
||||
|
||||
Demetrios:
|
||||
Maybe just, we'll start it off. We're going to get into it when it comes to what you're doing and really what the space looks like right now. Right. But I would love to hear a little bit of what you've been up to since, for the last two years, because I haven't talked to you.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, that's right. Well, several different things, actually. Right after we chatted last time, I joined Microsoft to work in developer relations. Microsoft has a big group of folks working in developer relations. And basically, for me, it signaled my shift away from regular software engineering. I was primarily doing software engineering and thought that perhaps with the books and some of the courses that I had published, it was time for me to get into more teaching and providing useful content, which is really something very rewarding. And in developer relations, in advocacy in general, it's kind of like a way of teaching. We demonstrate technology, how it works from a technical point of view.
|
||||
|
||||
Alfredo Deza:
|
||||
So aside from that, started working really closely with several different universities. I work with Georgia Tech, Oxford University, Carnegie Mellon University, and Duke University, where I've been working as an adjunct professor for a couple of years as well. So at Duke, what I do is I teach a couple of classes a year. One is on machine learning. Last year was machine learning operations, and this year it's going to, I think, hopefully I'm not messing anything up. I think we're going to shift a little bit to doing operations with large language models. And in the fall I teach a programming class for graduate students that want to join one of the graduate programs and they want to get a primer on Python. So I teach a little bit of that.
|
||||
|
||||
Alfredo Deza:
|
||||
And in the meantime, also in partnership with Duke, getting a lot of courses out on Coursera, and from large language models to doing stuff with Azure, to machine learning operations, to rust, I've been doing a lot of rust lately, which I really like. So, yeah, so a lot of different things, but I think the core pillar for me remains being able to teach and spread the knowledge.
|
||||
|
||||
Demetrios:
|
||||
Love it, man. And I know you've been diving into vector databases. Can you tell us more?
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, well, the thing is that when you're trying to teach, and yes, one of the courses that we had out for large language models was applying retrieval augmented generation, which is the basis for vector databases, to see how it works. This is how it works. These are the components that you need. Let's create an application from scratch and see how it works. And for those that don't know, retrieval augmented generation is kind of like having. The other day I saw a description about this, which I really like, which is a way of, it's kind of like having an open book test. So the large language model is the student, and they have an open book so they can see the answers and then repackage that into their own words and provide an answer, which is kind of like what we do with vector databases in the retrieval augmented generation pattern. We've been putting a lot of examples on how to do these, and in the case of Azure, you're enabling certain services.
|
||||
|
||||
Alfredo Deza:
|
||||
There's the Azure AI search service, which is really good. But sometimes when you're trying to teach specifically, it is useful to have a very straightforward way to do this and applying or creating a retrieval augmented generation pattern, it's kind of tricky, I think. We're not there yet to do it in a nice, straightforward way. So there are several different options, Qdrant being one of them. So usually I get asked, why are you using Qdrant? What's the big deal? Why are you picking these over all of the other ones? And to me it boils down to, aside from being renowned or recognized, that it works fairly well. There's one core component that is critical here, and that is it has to be very straightforward, very easy to set up so that I can teach it, because if it's easy, well, sort of like easy to or straightforward to teach, then you can take the next step and you can make it a little more complex, put other things around it, and that creates a great development experience and a learning experience as well. If something is very complex, if the list of requirements is very long, you're not going to be very happy, you're going to spend all this time trying to figure, and when you have, similar to what happens with automation, when you have a list of 20 different things that you need to, in order to, say, deploy a website, you're going to get things out of order, you're going to forget one thing, you're going to have a typo, you're going to mess it up, you're going to have to start from scratch, and you're going to get into a situation where you can't get out of it. And Qdrant does provide a very straightforward way to run the database, and that one is the in memory implementation with Python.
|
||||
|
||||
Alfredo Deza:
|
||||
So you can actually write a little bit of python once you install the libraries and say, I want to instantiate a vector database and I wanted to run it in memory. So for teaching, this is great. It's like, hey, of course it's not for production, but just write these couple of lines and let's get right into it. Let's just start populating these and see how it works. And it works. It's great. You don't need to have all of these, like, wow, let's launch Kubernetes over here and let's have all of these dynamic. No, why? I mean, sure, you want to create a business model and you want to launch to production eventually, and you want to have all that running perfect.
|
||||
|
||||
Alfredo Deza:
|
||||
But for this setup, like for understanding how it works, for trying baby steps into understanding vector databases, this is perfect. My one requirement, or my one wish list item is to have that in memory thing for rust. That would be pretty sweet, because I think it'll make teaching rust and retrieval augmented generation with rust much easier. I wouldn't have to worry about bringing up containers or external services. So that's the deal with rust. And I'll tell you one last story about why I think specifically making it easy to get started with so that I can teach it, so that others can learn from it, is crucial. I would say almost 50 years ago, maybe a little bit more, my dad went to Italy to have a course on athletics. My dad was involved in sports and he was going through this, I think it was like a six month specialization on athletics.
|
||||
|
||||
Alfredo Deza:
|
||||
And he was in class and it had been recent that the high jump had transitioned from one style to the other. The previous style, the old style right now is the old style. It's kind of like, it was kind of like over the bar. It was kind of like a weird style. And it had recently transitioned to a thing called the Fosbury flop. This person, his last name is Dick Fosbury, invented the Fosbury flop. He said, no, I'm just going to go straight at it, then do a little curve and then jump over it. And then he did, and then he started winning everything.
|
||||
|
||||
Alfredo Deza:
|
||||
And everybody's like, what this guy? Well, first they thought he was crazy, and they thought that dismissive of what he was trying to do. And there were people that sticklers that wanted to stay with the older style, but then he started beating records and winning medals, and so people were like, well, is this a good thing? Let's try it out. So there was a whole. They were casting doubt. It's like, is this really the thing? Is this really what we should be doing? So one of the questions that my dad had to answer in this specialization he did in Italy was like, which style is better, it's the old style or the new style? And so my dad said, it's the new style. And they asked him, why is the new style better? And he didn't choose the path of answering the, well, because this guy just won the Olympics or he just did a record over here that at the end is meaningless. What he said was, it is the better style because it's easier to teach and it is 100% correct. When you're teaching high jump, it is much easier to teach the Fosbury flop than the other style.
|
||||
|
||||
Alfredo Deza:
|
||||
It is super hard. So you start seeing this parallel in teaching and learning where, but with this one, you have all of these world records and things are going great. Well, great. But is anybody going to try, are you going to have more people looking into it or are you going to have less? What is it that we're trying to do here? Right.
|
||||
|
||||
Demetrios:
|
||||
Not going to lie, I did not see how you were going to land the plane on coming from the high jump into the vector database space, but you did it gracefully. That was well done. So, basically, the easier it is to teach, the more people are going to be able to jump on board and the more people are going to be able to get value out of it.
|
||||
|
||||
Sabrina Aquino:
|
||||
I absolutely love it, by the way. It's a pleasure to meet you, Alfredo. And I was actually about to ask you. I love your background as an olympic athlete. Right. And I was wondering, do you make any connections or how do we interact this background with your current teaching and AI? And do you see any similarities or something coming from that approach into what you've applied?
|
||||
|
||||
Alfredo Deza:
|
||||
Well, you're bringing a great point. It's taken me a very long time to feel comfortable talking about my professional sports past. I don't want to feel like I'm overwhelming anyone or trying to be like a show off. So I usually try not to mention, although I'm feeling more comfortable mentioning my professional past. But the only situations where I think it's good to talk about it is when I feel like there's a small chance that I might get someone thinking about the possibilities of what they can actually do and what they can try. And things that are seemingly complex might be achievable. So you mentioned similarities, but I think there are a couple of things that happen when you're an athlete in any sport, really, that you're trying to or you're operating at the very highest level and there's several things that happen there. You have to be consistent.
|
||||
|
||||
Alfredo Deza:
|
||||
And it's something that I teach my kids as well. I have one of my kids, he's like, I did really a lot of exercise today and then for a week he doesn't do anything else. And he's like, now I'm going to do exercise again. And she's going to do 4 hours. And it's like, wait a second, wait a second. It's okay. You want to do it. This is great.
|
||||
|
||||
Alfredo Deza:
|
||||
But no intensity. You need to be consistent. Oh, dad, you don't let me work out and it's like, no work out. Good, I support you, but you have to be consistent and slowly start ramping up and slowly start getting better. And it happens a lot with learning. We are in an era that concepts and things are advancing so fast that things are getting obsolete even faster. So you're always in this motion of trying to learn. So what I would say is the similarities are in the consistency.
|
||||
|
||||
Alfredo Deza:
|
||||
You have to keep learning, you have to keep applying yourself. But it can be like, oh, today I'm going to read this whole book from start to end and you're just going to learn everything about, I don't know, rust. It's like, well, no, try applying rust a little bit every day and feel comfortable with it. And at the very end you will do better. Like, you can't go with high intensity because you're going to get burned out, you're going to overwhelmed and it's not going to work out. You don't go to the Olympics by working out for like a few months. Actually, a very long time ago, a reporter asked me, how many months have you been working out preparing for the Olympics? It's like, what do you mean with how many months? I've been training my whole life for this. What are we talking about?
|
||||
|
||||
Demetrios:
|
||||
We're not talking in months or years. We're talking in lifetimes, right?
|
||||
|
||||
Alfredo Deza:
|
||||
So you have to take it easy. You can't do that. And beyond that, consistency. Consistency goes hand in hand with discipline. I came to the US in 2006. I don't live like I was born in Peru and I came to the US with no degree. I didn't go to college. Well, I went to college for a few months and then I dropped out and I didn't have a career, I didn't have experience.
|
||||
|
||||
Alfredo Deza:
|
||||
I was just recently married. I have never worked in my life because I used to be a professional athlete. And the only thing that I decided to do was to do amazing work, apply myself and try to keep learning and never stop learning. In the back of my mind, it's like, oh, I have a tremendous knowledge gap that I need to fulfill by learning. And actually, I have tremendous respect and I'm incredibly grateful by all of the people that opened doors for me and gave me an opportunity, one of them being Noah Giff, which I co authored a few books with him and some of the courses. And he actually taught me to write Python. I didn't know how to program. And he said, you know what? I think you should learn to write some python.
|
||||
|
||||
Alfredo Deza:
|
||||
And I was like, python? Why would I ever need to do that? And I did. He's like, let's just find something to automate. I mean, what a concept. Find something to apply automation. And every week on Fridays, we'll just take a look at it and that's it. And we did that for a while. And then he said, you know what? You should apply for speaking at Python. How can I be speaking at a conference when I just started learning? It's like your perspective is different.
|
||||
|
||||
Alfredo Deza:
|
||||
You just started learning these. You're going to do it in an interesting way. So I think those are concepts that are very important to me. Stay disciplined, stay consistent, and keep at it. The secret is that there's no secret. That's the bottom line. You have to keep consistent. Otherwise things are always making excuses.
|
||||
|
||||
Alfredo Deza:
|
||||
Is very simple.
|
||||
|
||||
Demetrios:
|
||||
The secret is there is no secret. That is beautiful. So you did kind of sprinkle this idea of, oh, I wish there was more stuff happening with Qdrant and rust. Can you talk a little bit more to that? Because one piece of Qdrant that people tend to love is that it's built in rust. Right. But also, I know that you mentioned before, could we get a little bit of this action so that I don't have to deal with any. What was it you were saying? The containers.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah. Right. Now, if you want to have a proof of concept, and I always go for like, what's the easiest, the most straightforward, the less annoying things I need to do, the better. And with Python, the Python API for Qdrant, you can just write a few lines and say, I want to create an instance in memory and then that's it. The database is created for you. This is very similar, or I would say actually almost identical to how you run SQLite. Sqlite is the embedded database you can create in memory. And it's actually how I teach SQL as well.
|
||||
|
||||
Alfredo Deza:
|
||||
When I have to teach SQl, I use sqlite. I think it's perfect. But in rust, like you said, Qdrant's backend is built on rust. There is no in memory implementation. So you are required to have an actual instance of the Qdrant database running. So you have a couple of options, but one of them probably means you'll have to bring up a container with Qdrant running and then you'll have to connect to that instance. So when you're teaching, the development environments are kind of constrained. Either you are in a lab somewhere like Crusader has labs, but those are self contained.
|
||||
|
||||
Alfredo Deza:
|
||||
It's kind of tricky to get them running 100%. You can run multiple containers at the same time. So things start becoming more complex. Not only more complex for the learner, but also in this case, like the teacher, me who wants to figure out how to make this all run in a very constrained environment. And that makes it tricky. And I fasted the team, by the way, and I was told that maybe at some point they can do some magic and put the in memory implementation on the rust side of things, which I think it would be tremendous.
|
||||
|
||||
Sabrina Aquino:
|
||||
We're going to advocate for that on our side. We're also going to be asking for it. And I think this is really good too. It really makes it easier. Me as a student not long ago, I do see what you mean. It's quite hard to get it all working very fast in the time of a class that you don't have a lot of time and students can get. I don't know, it's quite complex. I do get what you mean.
|
||||
|
||||
Sabrina Aquino:
|
||||
And you also are working both on the tech industry and on academia, which I think is super interesting. And I always kind of feel like those two are a bit disconnected sometimes. And I was wondering what you think that how important is the collaboration of these two areas considering how fast the AI space is going to right now? And what are your thoughts?
|
||||
|
||||
Alfredo Deza:
|
||||
Well, I don't like generalizing, but I'm going to generalize right now. I would say most universities are several steps behind, and there's a lot of complexities involved in higher education specifically. Most importantly, these institutions tend to be fairly large, and with fairly large institutions, what do you get? Oh, you get the magical bureaucracy for anything you want to do. Something like, oh, well, you need to talk to that department that needs to authorize something, that needs to go to some other department, and it's like, I'm going to change the curriculum. It's like, no, you can't. What does that mean? I have actually had conversations with faculty in universities where they say, listen, curricula. Yeah, we get that. We need to update it, but we change curricula every five years.
|
||||
|
||||
Alfredo Deza:
|
||||
And so. See you in a while. It's been three years. We have two more years to go. See you in a couple of years. And that's detrimental to students now. I get it. Building curricula, it's very hard.
|
||||
|
||||
Alfredo Deza:
|
||||
It takes a lot of work for the faculty to put something together. So it is something that, from a faculty perspective, it's like they're not going to get paid more if they update the curriculum.
|
||||
|
||||
Demetrios:
|
||||
Right.
|
||||
|
||||
Alfredo Deza:
|
||||
And it's a massive amount of work now that, of course, comes to the detriment of the learner. The student will be under service because they will have to go through curricula that is fairly dated. Now, there are situations and there are programs where this doesn't happen. And Duke, I've worked with several. They're teaching Llama file, which was built by Mozilla. And when did Llama file came out? It was just like a few months ago. And I think it's incredible. And I think those skills that are the ones that students need today in order to not only learn these things, but also be able to apply them when they're looking for a job or trying to professionally even apply them into their day to day, now that's one side of things.
|
||||
|
||||
Alfredo Deza:
|
||||
But there's the other aspect. In the case of Duke, as well as other universities out there, they're using these online platforms so that they can put courses out there faster. Do you really need to go through a four year program to understand how retrieval augmented generation works? Or how to implement it? I would argue no, but would you be better out, like, taking a course that will take you perhaps a couple of weeks to go through and be fairly proficient? I would say yes, 100%. And you see several institutions putting courses out there that are meaningful, that are useful, that they can cope with the speed at which things are needed. I think it's kind of good. And I think that sometimes we tend to think about knowledge and learning things, kind of like in a bubble, especially here in the US. I think there's this college is this magical place where all of the amazing things happen. And if you don't go to college, things are going to go very bad for you.
|
||||
|
||||
Alfredo Deza:
|
||||
And I don't think that's true. I think if you like college, if you like university, by all means take advantage of it. You want to experience it. That sounds great. I think there's tons of opportunity to do it outside of the university or the college setting and taking online courses from validated instructors. They have a good profile. Not someone that just dumped something on genetic AI and started.
|
||||
|
||||
Demetrios:
|
||||
Someone like you.
|
||||
|
||||
Alfredo Deza:
|
||||
Well, if you want to. Yeah, sure, why not? I mean, there's students that really like my teaching style. I think that's great. If you don't like my teaching style. Sometimes I tend to go a little bit slower because I don't want to overwhelm anyone. That's all good. But there is opportunity. And when I mention these things, people are like, oh, really? I'm not advertising for Coursera or anything else, but some of these platforms, if you pay a monthly fee, I think it's between $40 and $60.
|
||||
|
||||
Alfredo Deza:
|
||||
I think on the expensive side, you can take advantage of all of these courses and as much as you can take them. Sometimes even companies say, hey, you have a paid subscription, go take it all. And I've met people like that. It's like, this is incredible. I'm learning so much. Perfect. I think there's a mix of things. I don't think there's like a binary answer, like, oh, you need to do this, or, no, don't do that, and everything's going to be well again.
|
||||
|
||||
Demetrios:
|
||||
Yeah. Can you talk a little bit more about your course? And if I wanted to go on Coursera, what can I expect from.
|
||||
|
||||
Alfredo Deza:
|
||||
You know, and again, I don't think as much as I like talking about my courses and the things that I do, I want to emphasize, like, if someone is watching this video or listening into what we're talking about, find something that is interesting to you and find a course that kind of delivers that thing, that sliver of interesting stuff, and then try it out. I think that's the best way. Don't get overwhelmed by. It's like, is this the right vector database that I should be learning? Is this instructor? It's like, no, try it out. What's going to happen? You don't like it when you're watching a bad video series or docuseries on Netflix or any streaming platform? Do you just like, I pay my $10 a month, so I'm going to muster through this whole 20 more episodes of this thing that I don't like. It's meaningless. It doesn't matter. Just move on.
|
||||
|
||||
Alfredo Deza:
|
||||
So having said that, on Coursera specifically with Duke University, we tend to put courses out there that are going to be used in our programs in the things that I teach. For example, we just released the large language models. Specialization and specialization is a grouping of between four and six courses. So in there we have doing large language models with Azure, for example, introduction to generative AI, having a very simple rag pattern with Qdrant. I also have examples on how to do it with Azure AI search, which I think is pretty cool as well. How to do it locally with Llama file, which I think is great. You can have all of these large language models running locally, and then you have a little bit of Qdrant sprinkle over there, and then you have rack pattern. Now, I tend to teach with things that I really like, and I'll give you a quick example.
|
||||
|
||||
Alfredo Deza:
|
||||
I think there's three data sets that are one of the top three most used data sets in all of machine learning and data science. Those are the Boston housing market, the diabetes data set in the US, and the other one is the Titanic. And everybody uses those. And I don't really understand why. I mean, perhaps I do understand why. It's because they're easy, they're clean, they're ready to go. Nothing's ever wrong with these, and everybody has used them to boredom. But for the life of me, you wouldn't be able to convince me to use any of those, because these are not topics that I really care about and they don't resonate with me.
|
||||
|
||||
Alfredo Deza:
|
||||
The Titanic specifically is just horrid. Well, if I was 37 and I'm on first class and I'm male, would I survive? It's like, what are we trying to do here? How is this useful to anyone? So I tend to use things that I like, and I'm really passionate about wine. So I built my own data set, which is a collection of wines from all over the world, they have the ratings, they have the region, they have the type of grape and the notes and the name of the wine. So when I'm teaching them, like, look at this, this is amazing. It's wines from all over the world. So let's do a little bit of things here. So, for rag, what I was able to do is actually in the courses as well. I do, ah, I really know wines from Argentina, but these wines, it would be amazing if you can find me not a Malbec, but perhaps a cabernet franc.
|
||||
|
||||
Alfredo Deza:
|
||||
That is amazing. From, it goes through Qdrant, goes back to llama file using some large language model or even small language model, like the Phi 2 from Microsoft, I think is really good. And he goes, it tells. Yeah, sure. I get that you want to have some good wines. Here's some good stuff that I can give you. And so it's great, right? I think it's great. So I think those kinds of things that are interesting to the person that is teaching or presenting, I think that's the key, because whenever you're talking about things that are very boring, that you do not care about, things are not going to go well for you.
|
||||
|
||||
Alfredo Deza:
|
||||
I mean, if I didn't like teaching, if I didn't like vector databases, you would tell right away. It's like, well, yes, I've been doing stuff with the vector databases. They're good. Yeah, Qdrant, very good. You would tell right away. I can't lie. Very good.
|
||||
|
||||
Demetrios:
|
||||
You can't fool anybody.
|
||||
|
||||
Alfredo Deza:
|
||||
No.
|
||||
|
||||
Demetrios:
|
||||
Well, dude, this is awesome. We will drop a link to the chat. We will drop a link to the course in the chat so that in case anybody does want to go on this wine tasting journey with you, they can. And I'm sure there's all kinds of things that will spark the creativity of the students as they go through it, because when you were talking about that, I was like, oh, it would be really cool to make that same type of thing, but with ski resorts there, you go around the world. And if I want this type of ski resort, I'm going to just ask my chat bot. So I'm excited to see what people create with it. I also really appreciate you coming on here, giving us your time and talking through all this. It's been a pleasure, as always, Alfredo.
|
||||
|
||||
Demetrios:
|
||||
Thank you so much.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, thank you. Thank you for having me. Always happy to chat with you. I think Qdrant is doing a very solid product. Hopefully, my wish list item of in memory in rust comes to fruition, but I get it. Sometimes there are other priorities. It's all good. Yeah.
|
||||
|
||||
Alfredo Deza:
|
||||
If anyone wants to connect with me, I'm always active on LinkedIn primarily. Always happy to connect with folks and talk about learning and improving and always being a better person.
|
||||
|
||||
Demetrios:
|
||||
Excellent. Well, we will sign off, and if anyone else out there wants to come on here and talk to us about vector databases, we're always happy to have you. Feel free to reach out. And remember, don't get lost in vector space, folks. We will see you on the next one.
|
||||
|
||||
Sabrina Aquino:
|
||||
Good night. Thank you so much.
|
||||
@@ -0,0 +1,264 @@
|
||||
---
|
||||
draft: false
|
||||
title: Teaching Vector Databases at Scale - Alfredo Deza | Vector Space Talks
|
||||
slug: teaching-vector-db-at-scale
|
||||
short_description: Alfredo Deza tackles AI teaching, the intersection of
|
||||
technology and academia, and the value of consistent learning.
|
||||
description: Alfredo Deza discusses the practicality of machine learning
|
||||
operations, highlighting how personal interest in topics like wine datasets
|
||||
enhances engagement, while reflecting on the synergies between his
|
||||
professional sportsman discipline and the persistent, straightforward approach
|
||||
required for effectively educating on vector databases and large language
|
||||
models.
|
||||
preview_image: /blog/from_cms/alfredo-deza-bp-cropped.png
|
||||
date: 2024-04-09T03:06:00.000Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
- Vector Search
|
||||
- Retrieval Augmented Generation
|
||||
- Vector Space Talks
|
||||
- Coursera
|
||||
---
|
||||
> *"So usually I get asked, why are you using Qdrant? What's the big deal? Why are you picking these over all of the other ones? And to me it boils down to, aside from being renowned or recognized, that it works fairly well. There's one core component that is critical here, and that is it has to be very straightforward, very easy to set up so that I can teach it, because if it's easy, well, sort of like easy to or straightforward to teach, then you can take the next step and you can make it a little more complex, put other things around it, and that creates a great development experience and a learning experience as well.”*\
|
||||
— Alfredo Deza
|
||||
>
|
||||
|
||||
Alfredo is a software engineer, speaker, author, and former Olympic athlete working in Developer Relations at Microsoft. He has written several books about programming languages and artificial intelligence and has created online courses about the cloud and machine learning.
|
||||
|
||||
He currently is an Adjunct Professor at Duke University, and as part of his role, works closely with universities around the world like Georgia Tech, Duke University, Carnegie Mellon, and Oxford University where he often gives guest lectures about technology.
|
||||
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/4HFSrTJWxl7IgQj8j6kwXN?si=99H-p0fKQ0WuVEBJI9ugUw), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/3l6F6A_It0Q?feature=shared).***
|
||||
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/3l6F6A_It0Q?si=cFZGAh7995iHilcY" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
<iframe src="https://podcasters.spotify.com/pod/show/qdrant-vector-space-talk/embed/episodes/Teaching-Vector-Databases-at-Scale---Alfredo-Deza--Vector-Space-Talks-019-e2hhjlo/a-ab3qp7u" height="102px" width="400px" frameborder="0" scrolling="no"></iframe>
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
How does a former athlete such as Alfredo Deza end up in this AI and Machine Learning industry? That’s what we’ll find out in this episode of Vector Space Talks. Let’s understand how his background as an olympian offers a unique perspective on consistency and discipline that's a real game-changer in this industry.
|
||||
|
||||
Here are some things you’ll discover from this episode:
|
||||
|
||||
1. **The Intersection of Teaching and Tech:** Alfredo discusses on how to effectively bridge the gap between technical concepts and student understanding, especially when dealing with complex topics like vector databases.
|
||||
2. **Simplified Learning:** Dive into Alfredo's advocacy for simplicity in teaching methods, mirroring his approach with Qdrant and the potential for a Rust in-memory implementation aimed at enhancing learning experiences.
|
||||
3. **Beyond the Titanic Dataset:** Discover why Alfredo prefers to teach with a wine dataset he developed himself, underscoring the importance of using engaging subject matter in education.
|
||||
4. **AI Learning Acceleration:** Alfredo discusses the struggle universities face to keep pace with AI advancements and how online platforms can offer a more up-to-date curriculum.
|
||||
5. **Consistency is Key:** Alfredo draws parallels between the discipline required in high-level athletics and the ongoing learning journey in AI, zeroing in on his mantra, “There is no secret” to staying consistent.
|
||||
|
||||
> Fun Fact: Alfredo tells the story of athlete Dick Fosbury's invention of the Fosbury Flop to highlight the significance of teaching simplicity.
|
||||
>
|
||||
|
||||
## Show notes:
|
||||
|
||||
00:00 Teaching machine learning, Python to graduate students.\
|
||||
06:03 Azure AI search service simplifies teaching, Qdrant facilitates learning.\
|
||||
10:49 Controversy over high jump style.\
|
||||
13:18 Embracing past for inspiration, emphasizing consistency.\
|
||||
15:43 Consistent learning and practice lead to success.\
|
||||
20:26 Teaching SQL uses SQLite, Rust has limitations.\
|
||||
25:21 Online platforms improve and speed up education.\
|
||||
29:24 Duke and Coursera offer specialized language courses.\
|
||||
31:21 Passion for wines, creating diverse dataset.\
|
||||
35:00 Encouragement for vector db discussion, wrap up.\
|
||||
|
||||
## More Quotes from Alfredo:
|
||||
|
||||
*"Qdrant makes it straightforward. We use it in-memory for my classes and I would love to see something similar setup in Rust to make teaching even easier.”*\
|
||||
— Alfredo Deza
|
||||
|
||||
*"Retrieval augmented generation is kind of like having an open book test. So the large language model is the student, and they have an open book so they can see the answers and then repackage that into their own words and provide an answer.”*\
|
||||
— Alfredo Deza
|
||||
|
||||
*"With Qdrant, I appreciate that the use of the Python API is so simple. It avoids the complexity that comes from having a back-end system like in Rust where you need an actual instance of the database running.”*\
|
||||
— Alfredo Deza
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
What is happening? Everyone, welcome back to another vector space talks. I am Demetrios, and I am joined today by good old Sabrina. Where you at, Sabrina? Hello?
|
||||
|
||||
Sabrina Aquino:
|
||||
Hello, Demetrios. I'm from Brazil. I'm in Brazil right now. I know that you are traveling currently.
|
||||
|
||||
Demetrios:
|
||||
Where are you? At Kubecon in Paris. And it has been magnificent. But I could not wait to join the session today because we've got Alfredo coming at us.
|
||||
|
||||
Alfredo Deza:
|
||||
What's up, dude? Hi. How are you?
|
||||
|
||||
Demetrios:
|
||||
I'm good, man. It's been a while. I think the last time that we chatted was two years ago, maybe right before your book came out. When did the book come out?
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, something like that. I would say a couple of years ago. Yeah. I wrote, co authored practical machine learning operations with no gift. And it was published on O'Reilly.
|
||||
|
||||
Demetrios:
|
||||
Yeah. And that was, I think, two years ago. So you've been doing a lot of stuff since then. Let's be honest, you are maybe one of the most active men on the Internet. I always love seeing what you're doing. You're bringing immense value to everything that you touch. I'm really excited to be able to chat with you for this next 30 minutes.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, of course.
|
||||
|
||||
Demetrios:
|
||||
Maybe just, we'll start it off. We're going to get into it when it comes to what you're doing and really what the space looks like right now. Right. But I would love to hear a little bit of what you've been up to since, for the last two years, because I haven't talked to you.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, that's right. Well, several different things, actually. Right after we chatted last time, I joined Microsoft to work in developer relations. Microsoft has a big group of folks working in developer relations. And basically, for me, it signaled my shift away from regular software engineering. I was primarily doing software engineering and thought that perhaps with the books and some of the courses that I had published, it was time for me to get into more teaching and providing useful content, which is really something very rewarding. And in developer relations, in advocacy in general, it's kind of like a way of teaching. We demonstrate technology, how it works from a technical point of view.
|
||||
|
||||
Alfredo Deza:
|
||||
So aside from that, started working really closely with several different universities. I work with Georgia Tech, Oxford University, Carnegie Mellon University, and Duke University, where I've been working as an adjunct professor for a couple of years as well. So at Duke, what I do is I teach a couple of classes a year. One is on machine learning. Last year was machine learning operations, and this year it's going to, I think, hopefully I'm not messing anything up. I think we're going to shift a little bit to doing operations with large language models. And in the fall I teach a programming class for graduate students that want to join one of the graduate programs and they want to get a primer on Python. So I teach a little bit of that.
|
||||
|
||||
Alfredo Deza:
|
||||
And in the meantime, also in partnership with Duke, getting a lot of courses out on Coursera, and from large language models to doing stuff with Azure, to machine learning operations, to rust, I've been doing a lot of rust lately, which I really like. So, yeah, so a lot of different things, but I think the core pillar for me remains being able to teach and spread the knowledge.
|
||||
|
||||
Demetrios:
|
||||
Love it, man. And I know you've been diving into vector databases. Can you tell us more?
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, well, the thing is that when you're trying to teach, and yes, one of the courses that we had out for large language models was applying retrieval augmented generation, which is the basis for vector databases, to see how it works. This is how it works. These are the components that you need. Let's create an application from scratch and see how it works. And for those that don't know, retrieval augmented generation is kind of like having. The other day I saw a description about this, which I really like, which is a way of, it's kind of like having an open book test. So the large language model is the student, and they have an open book so they can see the answers and then repackage that into their own words and provide an answer, which is kind of like what we do with vector databases in the retrieval augmented generation pattern. We've been putting a lot of examples on how to do these, and in the case of Azure, you're enabling certain services.
|
||||
|
||||
Alfredo Deza:
|
||||
There's the Azure AI search service, which is really good. But sometimes when you're trying to teach specifically, it is useful to have a very straightforward way to do this and applying or creating a retrieval augmented generation pattern, it's kind of tricky, I think. We're not there yet to do it in a nice, straightforward way. So there are several different options, Qdrant being one of them. So usually I get asked, why are you using Qdrant? What's the big deal? Why are you picking these over all of the other ones? And to me it boils down to, aside from being renowned or recognized, that it works fairly well. There's one core component that is critical here, and that is it has to be very straightforward, very easy to set up so that I can teach it, because if it's easy, well, sort of like easy to or straightforward to teach, then you can take the next step and you can make it a little more complex, put other things around it, and that creates a great development experience and a learning experience as well. If something is very complex, if the list of requirements is very long, you're not going to be very happy, you're going to spend all this time trying to figure, and when you have, similar to what happens with automation, when you have a list of 20 different things that you need to, in order to, say, deploy a website, you're going to get things out of order, you're going to forget one thing, you're going to have a typo, you're going to mess it up, you're going to have to start from scratch, and you're going to get into a situation where you can't get out of it. And Qdrant does provide a very straightforward way to run the database, and that one is the in memory implementation with Python.
|
||||
|
||||
Alfredo Deza:
|
||||
So you can actually write a little bit of python once you install the libraries and say, I want to instantiate a vector database and I wanted to run it in memory. So for teaching, this is great. It's like, hey, of course it's not for production, but just write these couple of lines and let's get right into it. Let's just start populating these and see how it works. And it works. It's great. You don't need to have all of these, like, wow, let's launch Kubernetes over here and let's have all of these dynamic. No, why? I mean, sure, you want to create a business model and you want to launch to production eventually, and you want to have all that running perfect.
|
||||
|
||||
Alfredo Deza:
|
||||
But for this setup, like for understanding how it works, for trying baby steps into understanding vector databases, this is perfect. My one requirement, or my one wish list item is to have that in memory thing for rust. That would be pretty sweet, because I think it'll make teaching rust and retrieval augmented generation with rust much easier. I wouldn't have to worry about bringing up containers or external services. So that's the deal with rust. And I'll tell you one last story about why I think specifically making it easy to get started with so that I can teach it, so that others can learn from it, is crucial. I would say almost 50 years ago, maybe a little bit more, my dad went to Italy to have a course on athletics. My dad was involved in sports and he was going through this, I think it was like a six month specialization on athletics.
|
||||
|
||||
Alfredo Deza:
|
||||
And he was in class and it had been recent that the high jump had transitioned from one style to the other. The previous style, the old style right now is the old style. It's kind of like, it was kind of like over the bar. It was kind of like a weird style. And it had recently transitioned to a thing called the Fosbury flop. This person, his last name is Dick Fosbury, invented the Fosbury flop. He said, no, I'm just going to go straight at it, then do a little curve and then jump over it. And then he did, and then he started winning everything.
|
||||
|
||||
Alfredo Deza:
|
||||
And everybody's like, what this guy? Well, first they thought he was crazy, and they thought that dismissive of what he was trying to do. And there were people that sticklers that wanted to stay with the older style, but then he started beating records and winning medals, and so people were like, well, is this a good thing? Let's try it out. So there was a whole. They were casting doubt. It's like, is this really the thing? Is this really what we should be doing? So one of the questions that my dad had to answer in this specialization he did in Italy was like, which style is better, it's the old style or the new style? And so my dad said, it's the new style. And they asked him, why is the new style better? And he didn't choose the path of answering the, well, because this guy just won the Olympics or he just did a record over here that at the end is meaningless. What he said was, it is the better style because it's easier to teach and it is 100% correct. When you're teaching high jump, it is much easier to teach the Fosbury flop than the other style.
|
||||
|
||||
Alfredo Deza:
|
||||
It is super hard. So you start seeing this parallel in teaching and learning where, but with this one, you have all of these world records and things are going great. Well, great. But is anybody going to try, are you going to have more people looking into it or are you going to have less? What is it that we're trying to do here? Right.
|
||||
|
||||
Demetrios:
|
||||
Not going to lie, I did not see how you were going to land the plane on coming from the high jump into the vector database space, but you did it gracefully. That was well done. So, basically, the easier it is to teach, the more people are going to be able to jump on board and the more people are going to be able to get value out of it.
|
||||
|
||||
Sabrina Aquino:
|
||||
I absolutely love it, by the way. It's a pleasure to meet you, Alfredo. And I was actually about to ask you. I love your background as an olympic athlete. Right. And I was wondering, do you make any connections or how do we interact this background with your current teaching and AI? And do you see any similarities or something coming from that approach into what you've applied?
|
||||
|
||||
Alfredo Deza:
|
||||
Well, you're bringing a great point. It's taken me a very long time to feel comfortable talking about my professional sports past. I don't want to feel like I'm overwhelming anyone or trying to be like a show off. So I usually try not to mention, although I'm feeling more comfortable mentioning my professional past. But the only situations where I think it's good to talk about it is when I feel like there's a small chance that I might get someone thinking about the possibilities of what they can actually do and what they can try. And things that are seemingly complex might be achievable. So you mentioned similarities, but I think there are a couple of things that happen when you're an athlete in any sport, really, that you're trying to or you're operating at the very highest level and there's several things that happen there. You have to be consistent.
|
||||
|
||||
Alfredo Deza:
|
||||
And it's something that I teach my kids as well. I have one of my kids, he's like, I did really a lot of exercise today and then for a week he doesn't do anything else. And he's like, now I'm going to do exercise again. And she's going to do 4 hours. And it's like, wait a second, wait a second. It's okay. You want to do it. This is great.
|
||||
|
||||
Alfredo Deza:
|
||||
But no intensity. You need to be consistent. Oh, dad, you don't let me work out and it's like, no work out. Good, I support you, but you have to be consistent and slowly start ramping up and slowly start getting better. And it happens a lot with learning. We are in an era that concepts and things are advancing so fast that things are getting obsolete even faster. So you're always in this motion of trying to learn. So what I would say is the similarities are in the consistency.
|
||||
|
||||
Alfredo Deza:
|
||||
You have to keep learning, you have to keep applying yourself. But it can be like, oh, today I'm going to read this whole book from start to end and you're just going to learn everything about, I don't know, rust. It's like, well, no, try applying rust a little bit every day and feel comfortable with it. And at the very end you will do better. Like, you can't go with high intensity because you're going to get burned out, you're going to overwhelmed and it's not going to work out. You don't go to the Olympics by working out for like a few months. Actually, a very long time ago, a reporter asked me, how many months have you been working out preparing for the Olympics? It's like, what do you mean with how many months? I've been training my whole life for this. What are we talking about?
|
||||
|
||||
Demetrios:
|
||||
We're not talking in months or years. We're talking in lifetimes, right?
|
||||
|
||||
Alfredo Deza:
|
||||
So you have to take it easy. You can't do that. And beyond that, consistency. Consistency goes hand in hand with discipline. I came to the US in 2006. I don't live like I was born in Peru and I came to the US with no degree. I didn't go to college. Well, I went to college for a few months and then I dropped out and I didn't have a career, I didn't have experience.
|
||||
|
||||
Alfredo Deza:
|
||||
I was just recently married. I have never worked in my life because I used to be a professional athlete. And the only thing that I decided to do was to do amazing work, apply myself and try to keep learning and never stop learning. In the back of my mind, it's like, oh, I have a tremendous knowledge gap that I need to fulfill by learning. And actually, I have tremendous respect and I'm incredibly grateful by all of the people that opened doors for me and gave me an opportunity, one of them being Noah Giff, which I co authored a few books with him and some of the courses. And he actually taught me to write Python. I didn't know how to program. And he said, you know what? I think you should learn to write some python.
|
||||
|
||||
Alfredo Deza:
|
||||
And I was like, python? Why would I ever need to do that? And I did. He's like, let's just find something to automate. I mean, what a concept. Find something to apply automation. And every week on Fridays, we'll just take a look at it and that's it. And we did that for a while. And then he said, you know what? You should apply for speaking at Python. How can I be speaking at a conference when I just started learning? It's like your perspective is different.
|
||||
|
||||
Alfredo Deza:
|
||||
You just started learning these. You're going to do it in an interesting way. So I think those are concepts that are very important to me. Stay disciplined, stay consistent, and keep at it. The secret is that there's no secret. That's the bottom line. You have to keep consistent. Otherwise things are always making excuses.
|
||||
|
||||
Alfredo Deza:
|
||||
Is very simple.
|
||||
|
||||
Demetrios:
|
||||
The secret is there is no secret. That is beautiful. So you did kind of sprinkle this idea of, oh, I wish there was more stuff happening with Qdrant and rust. Can you talk a little bit more to that? Because one piece of Qdrant that people tend to love is that it's built in rust. Right. But also, I know that you mentioned before, could we get a little bit of this action so that I don't have to deal with any. What was it you were saying? The containers.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah. Right. Now, if you want to have a proof of concept, and I always go for like, what's the easiest, the most straightforward, the less annoying things I need to do, the better. And with Python, the Python API for Qdrant, you can just write a few lines and say, I want to create an instance in memory and then that's it. The database is created for you. This is very similar, or I would say actually almost identical to how you run SQLite. Sqlite is the embedded database you can create in memory. And it's actually how I teach SQL as well.
|
||||
|
||||
Alfredo Deza:
|
||||
When I have to teach SQl, I use sqlite. I think it's perfect. But in rust, like you said, Qdrant's backend is built on rust. There is no in memory implementation. So you are required to have an actual instance of the Qdrant database running. So you have a couple of options, but one of them probably means you'll have to bring up a container with Qdrant running and then you'll have to connect to that instance. So when you're teaching, the development environments are kind of constrained. Either you are in a lab somewhere like Crusader has labs, but those are self contained.
|
||||
|
||||
Alfredo Deza:
|
||||
It's kind of tricky to get them running 100%. You can run multiple containers at the same time. So things start becoming more complex. Not only more complex for the learner, but also in this case, like the teacher, me who wants to figure out how to make this all run in a very constrained environment. And that makes it tricky. And I fasted the team, by the way, and I was told that maybe at some point they can do some magic and put the in memory implementation on the rust side of things, which I think it would be tremendous.
|
||||
|
||||
Sabrina Aquino:
|
||||
We're going to advocate for that on our side. We're also going to be asking for it. And I think this is really good too. It really makes it easier. Me as a student not long ago, I do see what you mean. It's quite hard to get it all working very fast in the time of a class that you don't have a lot of time and students can get. I don't know, it's quite complex. I do get what you mean.
|
||||
|
||||
Sabrina Aquino:
|
||||
And you also are working both on the tech industry and on academia, which I think is super interesting. And I always kind of feel like those two are a bit disconnected sometimes. And I was wondering what you think that how important is the collaboration of these two areas considering how fast the AI space is going to right now? And what are your thoughts?
|
||||
|
||||
Alfredo Deza:
|
||||
Well, I don't like generalizing, but I'm going to generalize right now. I would say most universities are several steps behind, and there's a lot of complexities involved in higher education specifically. Most importantly, these institutions tend to be fairly large, and with fairly large institutions, what do you get? Oh, you get the magical bureaucracy for anything you want to do. Something like, oh, well, you need to talk to that department that needs to authorize something, that needs to go to some other department, and it's like, I'm going to change the curriculum. It's like, no, you can't. What does that mean? I have actually had conversations with faculty in universities where they say, listen, curricula. Yeah, we get that. We need to update it, but we change curricula every five years.
|
||||
|
||||
Alfredo Deza:
|
||||
And so. See you in a while. It's been three years. We have two more years to go. See you in a couple of years. And that's detrimental to students now. I get it. Building curricula, it's very hard.
|
||||
|
||||
Alfredo Deza:
|
||||
It takes a lot of work for the faculty to put something together. So it is something that, from a faculty perspective, it's like they're not going to get paid more if they update the curriculum.
|
||||
|
||||
Demetrios:
|
||||
Right.
|
||||
|
||||
Alfredo Deza:
|
||||
And it's a massive amount of work now that, of course, comes to the detriment of the learner. The student will be under service because they will have to go through curricula that is fairly dated. Now, there are situations and there are programs where this doesn't happen. And Duke, I've worked with several. They're teaching Llama file, which was built by Mozilla. And when did Llama file came out? It was just like a few months ago. And I think it's incredible. And I think those skills that are the ones that students need today in order to not only learn these things, but also be able to apply them when they're looking for a job or trying to professionally even apply them into their day to day, now that's one side of things.
|
||||
|
||||
Alfredo Deza:
|
||||
But there's the other aspect. In the case of Duke, as well as other universities out there, they're using these online platforms so that they can put courses out there faster. Do you really need to go through a four year program to understand how retrieval augmented generation works? Or how to implement it? I would argue no, but would you be better out, like, taking a course that will take you perhaps a couple of weeks to go through and be fairly proficient? I would say yes, 100%. And you see several institutions putting courses out there that are meaningful, that are useful, that they can cope with the speed at which things are needed. I think it's kind of good. And I think that sometimes we tend to think about knowledge and learning things, kind of like in a bubble, especially here in the US. I think there's this college is this magical place where all of the amazing things happen. And if you don't go to college, things are going to go very bad for you.
|
||||
|
||||
Alfredo Deza:
|
||||
And I don't think that's true. I think if you like college, if you like university, by all means take advantage of it. You want to experience it. That sounds great. I think there's tons of opportunity to do it outside of the university or the college setting and taking online courses from validated instructors. They have a good profile. Not someone that just dumped something on genetic AI and started.
|
||||
|
||||
Demetrios:
|
||||
Someone like you.
|
||||
|
||||
Alfredo Deza:
|
||||
Well, if you want to. Yeah, sure, why not? I mean, there's students that really like my teaching style. I think that's great. If you don't like my teaching style. Sometimes I tend to go a little bit slower because I don't want to overwhelm anyone. That's all good. But there is opportunity. And when I mention these things, people are like, oh, really? I'm not advertising for Coursera or anything else, but some of these platforms, if you pay a monthly fee, I think it's between $40 and $60.
|
||||
|
||||
Alfredo Deza:
|
||||
I think on the expensive side, you can take advantage of all of these courses and as much as you can take them. Sometimes even companies say, hey, you have a paid subscription, go take it all. And I've met people like that. It's like, this is incredible. I'm learning so much. Perfect. I think there's a mix of things. I don't think there's like a binary answer, like, oh, you need to do this, or, no, don't do that, and everything's going to be well again.
|
||||
|
||||
Demetrios:
|
||||
Yeah. Can you talk a little bit more about your course? And if I wanted to go on Coursera, what can I expect from.
|
||||
|
||||
Alfredo Deza:
|
||||
You know, and again, I don't think as much as I like talking about my courses and the things that I do, I want to emphasize, like, if someone is watching this video or listening into what we're talking about, find something that is interesting to you and find a course that kind of delivers that thing, that sliver of interesting stuff, and then try it out. I think that's the best way. Don't get overwhelmed by. It's like, is this the right vector database that I should be learning? Is this instructor? It's like, no, try it out. What's going to happen? You don't like it when you're watching a bad video series or docuseries on Netflix or any streaming platform? Do you just like, I pay my $10 a month, so I'm going to muster through this whole 20 more episodes of this thing that I don't like. It's meaningless. It doesn't matter. Just move on.
|
||||
|
||||
Alfredo Deza:
|
||||
So having said that, on Coursera specifically with Duke University, we tend to put courses out there that are going to be used in our programs in the things that I teach. For example, we just released the large language models. Specialization and specialization is a grouping of between four and six courses. So in there we have doing large language models with Azure, for example, introduction to generative AI, having a very simple rag pattern with Qdrant. I also have examples on how to do it with Azure AI search, which I think is pretty cool as well. How to do it locally with Llama file, which I think is great. You can have all of these large language models running locally, and then you have a little bit of Qdrant sprinkle over there, and then you have rack pattern. Now, I tend to teach with things that I really like, and I'll give you a quick example.
|
||||
|
||||
Alfredo Deza:
|
||||
I think there's three data sets that are one of the top three most used data sets in all of machine learning and data science. Those are the Boston housing market, the diabetes data set in the US, and the other one is the Titanic. And everybody uses those. And I don't really understand why. I mean, perhaps I do understand why. It's because they're easy, they're clean, they're ready to go. Nothing's ever wrong with these, and everybody has used them to boredom. But for the life of me, you wouldn't be able to convince me to use any of those, because these are not topics that I really care about and they don't resonate with me.
|
||||
|
||||
Alfredo Deza:
|
||||
The Titanic specifically is just horrid. Well, if I was 37 and I'm on first class and I'm male, would I survive? It's like, what are we trying to do here? How is this useful to anyone? So I tend to use things that I like, and I'm really passionate about wine. So I built my own data set, which is a collection of wines from all over the world, they have the ratings, they have the region, they have the type of grape and the notes and the name of the wine. So when I'm teaching them, like, look at this, this is amazing. It's wines from all over the world. So let's do a little bit of things here. So, for rag, what I was able to do is actually in the courses as well. I do, ah, I really know wines from Argentina, but these wines, it would be amazing if you can find me not a Malbec, but perhaps a cabernet franc.
|
||||
|
||||
Alfredo Deza:
|
||||
That is amazing. From, it goes through Qdrant, goes back to llama file using some large language model or even small language model, like the Phi 2 from Microsoft, I think is really good. And he goes, it tells. Yeah, sure. I get that you want to have some good wines. Here's some good stuff that I can give you. And so it's great, right? I think it's great. So I think those kinds of things that are interesting to the person that is teaching or presenting, I think that's the key, because whenever you're talking about things that are very boring, that you do not care about, things are not going to go well for you.
|
||||
|
||||
Alfredo Deza:
|
||||
I mean, if I didn't like teaching, if I didn't like vector databases, you would tell right away. It's like, well, yes, I've been doing stuff with the vector databases. They're good. Yeah, Qdrant, very good. You would tell right away. I can't lie. Very good.
|
||||
|
||||
Demetrios:
|
||||
You can't fool anybody.
|
||||
|
||||
Alfredo Deza:
|
||||
No.
|
||||
|
||||
Demetrios:
|
||||
Well, dude, this is awesome. We will drop a link to the chat. We will drop a link to the course in the chat so that in case anybody does want to go on this wine tasting journey with you, they can. And I'm sure there's all kinds of things that will spark the creativity of the students as they go through it, because when you were talking about that, I was like, oh, it would be really cool to make that same type of thing, but with ski resorts there, you go around the world. And if I want this type of ski resort, I'm going to just ask my chat bot. So I'm excited to see what people create with it. I also really appreciate you coming on here, giving us your time and talking through all this. It's been a pleasure, as always, Alfredo.
|
||||
|
||||
Demetrios:
|
||||
Thank you so much.
|
||||
|
||||
Alfredo Deza:
|
||||
Yeah, thank you. Thank you for having me. Always happy to chat with you. I think Qdrant is doing a very solid product. Hopefully, my wish list item of in memory in rust comes to fruition, but I get it. Sometimes there are other priorities. It's all good. Yeah.
|
||||
|
||||
Alfredo Deza:
|
||||
If anyone wants to connect with me, I'm always active on LinkedIn primarily. Always happy to connect with folks and talk about learning and improving and always being a better person.
|
||||
|
||||
Demetrios:
|
||||
Excellent. Well, we will sign off, and if anyone else out there wants to come on here and talk to us about vector databases, we're always happy to have you. Feel free to reach out. And remember, don't get lost in vector space, folks. We will see you on the next one.
|
||||
|
||||
Sabrina Aquino:
|
||||
Good night. Thank you so much.
|
||||
@@ -1,16 +1,16 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: "VirtualBrain: Best RAG to unleash the real power of AI - Guillaume
|
||||
Marquis | Vector Space Talks"
|
||||
slug: virtualbrain-best-rag
|
||||
short_description: Delve into the complex world of information retrieval with
|
||||
Guillaume Marquis, CTO & Co-founder at VirtualBrain.
|
||||
description: Guillaume Marquis, CTO & Co-founder at VirtualBrain, reveals the
|
||||
short_description: Let's explore information retrieval with Guillaume Marquis,
|
||||
CTO & Co-Founder at VirtualBrain.
|
||||
description: Guillaume Marquis, CTO & Co-Founder at VirtualBrain, reveals the
|
||||
mechanics of advanced document retrieval with RAG technology, discussing the
|
||||
challenges of scalability, up-to-date information, and navigating user
|
||||
feedback to enhance the productivity of knowledge workers.
|
||||
preview_image: /blog/from_cms/guillaume-marquis-2-cropped.png
|
||||
date: 2024-03-11T08:35:54.200Z
|
||||
date: 2024-03-27T12:41:51.859Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -23,19 +23,19 @@ tags:
|
||||
— Guillaume Marquis
|
||||
>
|
||||
|
||||
Guillaume Marquis, a dedicated Engineer and AI enthusiast, serves as the Chief Technology Officer and Co-founder of VirtualBrain, an innovative AI company. He is committed to exploring novel approaches to integrating artificial intelligence into everyday life, driven by a passion for advancing the field and its applications.
|
||||
Guillaume Marquis, a dedicated Engineer and AI enthusiast, serves as the Chief Technology Officer and Co-Founder of VirtualBrain, an innovative AI company. He is committed to exploring novel approaches to integrating artificial intelligence into everyday life, driven by a passion for advancing the field and its applications.
|
||||
|
||||
***Listen to the episode on Spotify, Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on YouTube.***
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/20iFzv2sliYRSHRy1QHq6W?si=xZqW2dF5QxWsAN4nhjYGmA), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/v85HqNqLQcI?feature=shared).***
|
||||
|
||||
[embed YouTube video here]
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/v85HqNqLQcI?si=hjUiIhWxsDVO06-H" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
[embed anchor.fm podcast here]
|
||||
<iframe src="https://podcasters.spotify.com/pod/show/qdrant-vector-space-talk/embed/episodes/VirtualBrain-Best-RAG-to-unleash-the-real-power-of-AI---Guillaume-Marquis--Vector-Space-Talks-017-e2grbfg/a-ab22dgt" height="102px" width="400px" frameborder="0" scrolling="no"></iframe>
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
Who knew that document retrieval could be creative? Guillaume and VirtualBrain help draft sales proposals using past reports. Fascinating how tech aids deep work beyond basic search tasks.
|
||||
Who knew that document retrieval could be creative? Guillaume and VirtualBrain help draft sales proposals using past reports. It's fascinating how tech aids deep work beyond basic search tasks.
|
||||
|
||||
Delving into document retrieval and AI assistance, Guillaume furthermore unpacks the intricacies of sifting through vast data using a scoring system, the virtue of RAG for deep work, and combating the 'illusion of work', enhancing insights for knowledge workers while confronting the challenges of scalability and user feedback on hallucinations.
|
||||
Tackling document retrieval and AI assistance, Guillaume furthermore unpacks the ins and outs of searching through vast data using a scoring system, the virtue of RAG for deep work, and going through the 'illusion of work', enhancing insights for knowledge workers while confronting the challenges of scalability and user feedback on hallucinations.
|
||||
|
||||
Here are some key insight from this episode you need to look out for:
|
||||
|
||||
@@ -51,7 +51,7 @@ Here are some key insight from this episode you need to look out for:
|
||||
|
||||
## Show notes:
|
||||
|
||||
00:00 Hosts’ and guest recommendations.\
|
||||
00:00 Hosts and guest recommendations.\
|
||||
09:01 Leveraging past knowledge to create new proposals.\
|
||||
12:33 Ingesting and parsing documents for context retrieval.\
|
||||
14:26 Creating and storing data, performing advanced searches.\
|
||||
@@ -69,7 +69,7 @@ Here are some key insight from this episode you need to look out for:
|
||||
*"We only exclusively use open source tools because of security aspects and stuff like that. That's why also we are using Qdrant one of the important point on that. So we have a system, we are using this serverless stuff to ingest document over time.”*\
|
||||
— Guillaume Marquis
|
||||
|
||||
*"One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today Chat GPT answer really fast. So your users, even if the question is way more complicated to answer than a basic Chat GPT question, they want to have their answer in seconds. So you have also this challenge that is really you have to take care.”*\
|
||||
*"One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today ChatGPT answers really fast. So your users, even if the question is way more complicated to answer than a basic ChatGPT question, they want to have their answer in seconds. So you have also this challenge that you really have to take care.”*\
|
||||
— Guillaume Marquis
|
||||
|
||||
*"Our AI is not trained to write you a speech based on Shakespeare and with the style of Martin Luther King. It's not the purpose of the tool. So if you ask something that is out of the box, he will just say like, okay, I don't know how to answer that. And that's an important point. That's a feature by itself to be able to not go outside of the box.”*\
|
||||
@@ -80,7 +80,7 @@ Demetrios:
|
||||
So, dude, I'm excited for this talk. Before we get into it, I want to make sure that we have some pre conversation housekeeping items that go out, one of which being, as always, we're doing these vector space talks and everyone is encouraged and invited to join in. Ask your questions, let us know where you're calling in from, let us know what you're up to, what your use case is, and feel free to drop any questions that you may have in the chat. We will be monitoring it like a hawk. Today I am joined by none other than Sabrina. How are you doing, Sabrina?
|
||||
|
||||
Sabrina Aquino:
|
||||
What's up, dementias? I'm doing great. Excited to be here. I just love seeing what amazing stuff people are building with Qdrant and. Yeah, let's get into it.
|
||||
What's up, Demetrios? I'm doing great. Excited to be here. I just love seeing what amazing stuff people are building with Qdrant and. Yeah, let's get into it.
|
||||
|
||||
Demetrios:
|
||||
Yeah. So I think I see Sabrina's wearing a special shirt which is don't get lost in vector space shirt. If anybody wants a shirt like that. There we go. Well, we got you covered, dude. You will get one at your front door soon enough. If anybody else wants one, come on here. Present at the next vector space talks.
|
||||
@@ -89,7 +89,7 @@ Demetrios:
|
||||
We're excited to have you. And we've got one last thing that I think is fun that we can talk about before we jump into the tech piece of the conversation. And that is I told Sabrina to get ready with some recommendations. Know vector databases, they can be used occasionally for recommendation systems, but nothing's better than getting that hidden gem from your friend. And right now what we're going to try and do is give you a few hidden gems so that the next time the recommendation engine is working for you, it's working in your favor. And Sabrina, I asked you to give me one music that you can recommend, one show and one rando. So basically one random thing that you can recommend to us.
|
||||
|
||||
Sabrina Aquino:
|
||||
So I've picked. I thought about this. Okay, I give it some thought. The movie would be catch me if you can by Leo DiCaprio and tone Hanks. Have you guys watched it? Really good movie. The song would be oh, children by knee cave and the bad scenes. Also very good song. And the random recommendation is my favorite scented candle, which is citrus notes, sea salt and cedar.
|
||||
So I've picked. I thought about this. Okay, I give it some thought. The movie would be Catch Me If You Can by Leo DiCaprio and Tom Hanks. Have you guys watched it? Really good movie. The song would be oh, children by knee cave and the bad scenes. Also very good song. And the random recommendation is my favorite scented candle, which is citrus notes, sea salt and cedar.
|
||||
|
||||
Sabrina Aquino:
|
||||
So there you go.
|
||||
@@ -98,13 +98,13 @@ Demetrios:
|
||||
A scented candle as a recommendation. I like it. I think that's cool. I didn't exactly tell you to get ready with that. So I'll go next, then you can have some more time to think. So for anybody that's joining in, we're just giving a few recommendations to help your own recommendation engines at home. And we're going to get into this conversation about rags in just a moment. But my song is with.
|
||||
|
||||
Demetrios:
|
||||
Oh, my God. I've been listening to it because I didn't think that they had it on Spotify, but I found it this morning and I was so happy that they did. And it is Bill Evans and Chet Baker. Basically, their whole album, the legendary sessions, is just like, incredible. But the first song on that album is called alone together. And when Chet Baker starts playing his little trombone, my God, it is like you can feel emotion. You can touch it. That is what I would recommend.
|
||||
Oh, my God. I've been listening to it because I didn't think that they had it on Spotify, but I found it this morning and I was so happy that they did. And it is Bill Evans and Chet Baker. Basically, their whole album, the legendary sessions, is just like, incredible. But the first song on that album is called Alone Together. And when Chet Baker starts playing his little trombone, my God, it is like you can feel emotion. You can touch it. That is what I would recommend.
|
||||
|
||||
Demetrios:
|
||||
Anyone out there? I'll drop a link in the chat if you like it. The film or series. This fool, if you speak Spanish, it's even better. It is amazing series. Get that, do it. And as the rando thing, I've been having Rishi mushroom powder in my coffee in the mornings. I highly recommend it. All right, last one, let's get into your recommendations and then we'll get into this rag chat.
|
||||
|
||||
Guillaume Marquis:
|
||||
So, yeah, I sucked a little bit. So for the song, I think I will give something like, because I'm french, I think you can hear it. So I will choose get lucky of daft Punk and because I am a little bit sad of the end of their collaboration. So, yeah, just like, I cannot forget it. And it's a really good music. Like, miss them as a movie, maybe something like I really enjoy. So we have a lot of french movies that are really nice, but something more international maybe, and more mainstream. Jungle of Tarantino, that is really a good movie and really enjoy it.
|
||||
So, yeah, I sucked a little bit. So for the song, I think I will give something like, because I'm french, I think you can hear it. So I will choose Get Lucky of Daft Punk and because I am a little bit sad of the end of their collaboration. So, yeah, just like, I cannot forget it. And it's a really good music. Like, miss them as a movie, maybe something like I really enjoy. So we have a lot of french movies that are really nice, but something more international maybe, and more mainstream. Jungle of Tarantino, that is really a good movie and really enjoy it.
|
||||
|
||||
Guillaume Marquis:
|
||||
I watched it several times and still a good movie to watch. And random thing, maybe a city. A city to go to visit. I really enjoyed. It's hard to choose. Really hard to choose a place in general. Okay, Florence, like in Italy.
|
||||
@@ -140,10 +140,10 @@ Demetrios:
|
||||
I have the million dollar question that I think is probably coming through everyone's head is like, you're retrieving so many documents, how are you evaluating your retrieval?
|
||||
|
||||
Guillaume Marquis:
|
||||
That's definitely the $1 million question. It's a toss task to do, to be honest. To be fair. Currently what we are doing is that we monitor every tasks of the process, so we have the output of every tasks. On each tasks we use a scoring system to evaluate if it's relevant to the initial question or the initial task of the user. And we have a global scoring system on all the system. So it's quite od, it's a little bit empiric, but it works for now. And it really help us to also improve over time all the tasks and all the processes that are done by the tool.
|
||||
That's definitely the $1 million question. It's a toss task to do, to be honest. To be fair. Currently what we are doing is that we monitor every tasks of the process, so we have the output of every tasks. On each tasks we use a scoring system to evaluate if it's relevant to the initial question or the initial task of the user. And we have a global scoring system on all the system. So it's quite odd, it's a little bit empiric, but it works for now. And it really help us to also improve over time all the tasks and all the processes that are done by the tool.
|
||||
|
||||
Guillaume Marquis:
|
||||
So it's really important. And for instance, you have this kind of framework that is called ragtriad. That is a way to evaluate rag on the accuracy of the context you retrieve on the link with the initial question and so on, several parameters. And you can really have a first way to evaluate the quality of answers and the quality of everything on each steps.
|
||||
So it's really important. And for instance, you have this kind of framework that is called RAGtriad. That is a way to evaluate rag on the accuracy of the context you retrieve on the link with the initial question and so on, several parameters. And you can really have a first way to evaluate the quality of answers and the quality of everything on each steps.
|
||||
|
||||
Sabrina Aquino:
|
||||
I love it. Can you go more into the tech that you use for each one of these steps in architecture?
|
||||
@@ -158,7 +158,7 @@ Guillaume Marquis:
|
||||
So basically we are creating unbelieving, we are storing it into Qdrant. We are performing similarity search to retrieve documents based on title summary filtering, on tags, on the semantic context. And we have also some keyword search, but it's more for specific tasks, like when we know that we need a specific document, at some point we are searching it with a keyword search. So it's like a kind of ebrid system that is using deterministic approach with filtering with tags, and a probabilistic approach with selecting document with this ebot search, and doing a scoring system after that to get what is the most relevant document and to select how much content we will take from each document. It's a little bit techy, but it's really cool to create and we have a way to evolve it and to improve it.
|
||||
|
||||
Demetrios:
|
||||
That's what we like around here, man. We want the techie stuff. That's what I think everybody signed up for. So that's very cool. One question that definitely comes up a lot when it comes to rags and when you're ingesting documents, and then when you're retrieving documents and updating documents, how do you make sure that the documents that you are, let's say, I know there's probably a hypothetical HR scenario where the company has a certain policy and they say you can have european style holidays, you get like three months of holidays a year, or even french style holidays. Basically, you just don't work. And whenever you want, you can work, you don't work. And then all of a sudden a US company comes and takes it over and they say, no, you guys don't get holidays.
|
||||
That's what we like around here, man. We want the techie stuff. That's what I think everybody signed up for. So that's very cool. One question that definitely comes up a lot when it comes to rags and when you're ingesting documents, and then when you're retrieving documents and updating documents, how do you make sure that the documents that you are, let's say, I know there's probably a hypothetical HR scenario where the company has a certain policy and they say you can have European style holidays, you get like three months of holidays a year, or even French style holidays. Basically, you just don't work. And whenever you want, you can work, you don't work. And then all of a sudden a US company comes and takes it over and they say, no, you guys don't get holidays.
|
||||
|
||||
Demetrios:
|
||||
Even when you do get holidays, you're not working or you are working and so you have to update all the HR documents, right? So now when you have this knowledge worker that is creating something, or when you have anyone that is getting help, like this copilot help, how do you make sure that the information that person is getting is the most up to date information possible?
|
||||
@@ -170,7 +170,7 @@ Demetrios:
|
||||
I'm coming with the hits today. I don't know what you were looking for.
|
||||
|
||||
Guillaume Marquis:
|
||||
That's a really good question. So basically you have several possibilities on that. First one you have like this PowerPoint presentation, v one, v two, vf vf one, vf two, et cetera. That's a mess in the knowledge bases and sometimes you just want to use the most updated up to date documents. So basically we can filter on the created ad and the date of the documents. Sometimes you want to also compare the evolution of the process over time. So that's another use case. Basically we base.
|
||||
That's a really good question. So basically you have several possibilities on that. First one you have like this PowerPoint presentation. That's a mess in the knowledge bases and sometimes you just want to use the most updated up to date documents. So basically we can filter on the created ad and the date of the documents. Sometimes you want to also compare the evolution of the process over time. So that's another use case. Basically we base.
|
||||
|
||||
Guillaume Marquis:
|
||||
So during the ingestion we are analyzing if date is inside the document, because sometimes in documentation you have like the date at the end of the document or at the beginning of the document. That's a first way to do it. We have the date of the creation of the document, but it's not a source of truth because sometimes you created it after or you duplicated it and the date is not the same, depending if you are working on Windows, Microsoft, stuff like that. It's definitely a mess. And also we compare documents. So when we retry the documents and documents are really similar one to each other, we keep it in mind and we try to give more information as possible. Sometimes it's not possible, so it's not 100%, it's not bulletproof, but it's a real question of that. So it's a partial answer of your question, but it's like some way we are today filtering and answering on this special topic.
|
||||
@@ -185,13 +185,13 @@ Sabrina Aquino:
|
||||
Challenging.
|
||||
|
||||
Guillaume Marquis:
|
||||
One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today chgptpt answer really fast. So your users, even if the question is way more complicated to answer than a basic Chat GPT question, they want to have their answer in seconds. So you have also this challenge that is really you have to take care. So it's quite challenging and it's like this industrial supply chain. So when you upgrade something, you have to be sure that everything is working well on the other side. And that's a real challenge to handle.
|
||||
One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today Chat GPT answer really fast. So your users, even if the question is way more complicated to answer than a basic Chat GPT question, they want to have their answer in seconds. So you have also this challenge that is really you have to take care. So it's quite challenging and it's like this industrial supply chain. So when you upgrade something, you have to be sure that everything is working well on the other side. And that's a real challenge to handle.
|
||||
|
||||
Guillaume Marquis:
|
||||
And we are still on it because we are still evolving and getting more data. And at the end of the day, you have to be sure that everything is working well in terms of LLM, but in terms of research and in terms also a few weeks to give some insight to the user of what is working under the hood, to give them the possibility to wait a few seconds more, but starting to give them pieces of answer.
|
||||
|
||||
Demetrios:
|
||||
Yeah, it's funny you say that because I remember talking to somebody that was working@u.com and they were saying how there's like the actual time. So they were calling it something like perceived time and real, like actual time. So you as an end user, if you get asked a question or maybe there's like a trivia quiz while the question is coming up, then it seems like it's not actually taking as long as it is. Even if it takes 5 seconds, it's a little bit cooler. Or as you were mentioning, I remember reading some paper, I think, on how people are a lot less anxious if they see the words starting to pop up like that and they see like, okay, it's not just I'm waiting and then the whole answer gets spit back out at me. It's like I see the answer forming as it is in real time. And so that can calm people's nerves too.
|
||||
Yeah, it's funny you say that because I remember talking to somebody that was working at you.com and they were saying how there's like the actual time. So they were calling it something like perceived time and real, like actual time. So you as an end user, if you get asked a question or maybe there's like a trivia quiz while the question is coming up, then it seems like it's not actually taking as long as it is. Even if it takes 5 seconds, it's a little bit cooler. Or as you were mentioning, I remember reading some paper, I think, on how people are a lot less anxious if they see the words starting to pop up like that and they see like, okay, it's not just I'm waiting and then the whole answer gets spit back out at me. It's like I see the answer forming as it is in real time. And so that can calm people's nerves too.
|
||||
|
||||
Guillaume Marquis:
|
||||
Yeah, definitely. Human's brain is like marvelous on that. And you have a lot of stuff. Like, one of my favorites is the illusion of work. Do you know it? It's the total opposite. If you have something that seems difficult to do, adding more time of processing. So the user will imagine that it's really an OD task to do. And so that's really funny.
|
||||
@@ -200,7 +200,7 @@ Demetrios:
|
||||
So funny like that.
|
||||
|
||||
Guillaume Marquis:
|
||||
Yeah. Yes. It's the opposite of what you will think if you create a product, but that's real stuff. And sometimes just to output them that you are performing toss tasks in the background, it helps them to. Oh, yes. My question was really like a complex question, like you have a lot of work to do. It's Axx word like. If you answer too fast, they will not trust the answer.
|
||||
Yeah. Yes. It's the opposite of what you will think if you create a product, but that's real stuff. And sometimes just to output them that you are performing toss tasks in the background, it helps them to. Oh, yes. My question was really like a complex question, like you have a lot of work to do. It's Axe word like. If you answer too fast, they will not trust the answer.
|
||||
|
||||
Guillaume Marquis:
|
||||
And it's the opposite if you answer too slow. You can have this. Okay. But it should be dumb because it's really slow. So it's a dumb AI or stuff like that. So that's really funny. My co founder actually was a product guy, so really focused on product, and he really loves this kind of stuff.
|
||||
@@ -218,7 +218,7 @@ Demetrios:
|
||||
Some tell me more. Yeah.
|
||||
|
||||
Guillaume Marquis:
|
||||
So we tried the classic postgre page vectors, that is, I think we tried it like 30 minutes, and we realized really fast that it was really not good for our use case. We tried wavyt, we tried Milvus, we tried Qdrant, we tried a lot. We prefer use open source because of security issues. We tried Pinecone initially, we were on Pinecone at the beginning of the company. And so the most important point, so we have the speed of the tool, we have the scalability we have also, maybe it's a little bit dumb to say that, but we have also the API. I remember using Pinecone and trying just to get all vectors and it was not possible somehow, and you have this dumb stuff that are sometimes really strange. And if you have a tool that is 100% made for your use case with people that are working on it, really dedicated on that, and that are aligned with your vision of what is the evolution of this. I think it's like the best tool you have to choose.
|
||||
So we tried the classic postgres page vectors, that is, I think we tried it like 30 minutes, and we realized really fast that it was really not good for our use case. We tried Weaviate, we tried Milvus, we tried Qdrant, we tried a lot. We prefer use open source because of security issues. We tried Pinecone initially, we were on Pinecone at the beginning of the company. And so the most important point, so we have the speed of the tool, we have the scalability we have also, maybe it's a little bit dumb to say that, but we have also the API. I remember using Pinecone and trying just to get all vectors and it was not possible somehow, and you have this dumb stuff that are sometimes really strange. And if you have a tool that is 100% made for your use case with people that are working on it, really dedicated on that, and that are aligned with your vision of what is the evolution of this. I think it's like the best tool you have to choose.
|
||||
|
||||
Demetrios:
|
||||
So one thing that I would love to hear about too, is when you're looking at your system and you're looking at just the product in general, what are some of the key metrics that you are constantly monitoring, and how do you know that you're hitting them or you're not? And then if you're not hitting them, what are some ways that you debug the situation?
|
||||
@@ -290,7 +290,7 @@ Sabrina Aquino:
|
||||
Yeah, I think I'm just very interesting to know from a user perspective, from a virtual brain, how are traditional models worse or what kind of errors virtual brain fixes in their structure, that users find it better that way.
|
||||
|
||||
Guillaume Marquis:
|
||||
I think in this particular, so we talked about hallucinations, I think it's like one of the main issues people have on classic elements. We really think that when you create a one size fit all tool, you have some chole because you have to manage different approaches, like when you are creating copilot as Microsoft, you have to under the use cases of, and I really think so. Our AI is not trained to write you a speech based on Shakespeare and with the style of Martin Luther Luther King. It's not the purpose of the tool. So if you ask something that is out of the box, he will just say like, okay, I don't know how to answer that. And that's an important point. That's a feature by itself to be able to not go outside of the box. And so we did this choice of putting the AI inside the box, the box that is containing basically all the knowledge of your company, all the retrieved knowledge.
|
||||
I think in this particular, so we talked about hallucinations, I think it's like one of the main issues people have on classic elements. We really think that when you create a one size fit all tool, you have some chole because you have to manage different approaches, like when you are creating copilot as Microsoft, you have to under the use cases of, and I really think so. Our AI is not trained to write you a speech based on Shakespeare and with the style of Martin Luther King. It's not the purpose of the tool. So if you ask something that is out of the box, he will just say like, okay, I don't know how to answer that. And that's an important point. That's a feature by itself to be able to not go outside of the box. And so we did this choice of putting the AI inside the box, the box that is containing basically all the knowledge of your company, all the retrieved knowledge.
|
||||
|
||||
Guillaume Marquis:
|
||||
Actually we do not have a lot of hallucination, I will not say like 0%, but it's close to zero. Because we analyze a question, we put the AI in a box, we enforce the AI to think about the answer before answering, and we analyze also the answer to know if the answer is relevant. And that's an important point that we are fixing and we fix for our user and we prefer yes, to give like non answers and a bad answer.
|
||||
|
||||
@@ -63,10 +63,9 @@ curl -X PUT http://localhost:6333/collections/test_collection1 \
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -195,9 +194,8 @@ curl -X PUT http://localhost:6333/collections/test_collection2 \
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -325,10 +323,10 @@ curl -X PUT http://localhost:6333/collections/test_collection3 \
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -487,10 +485,9 @@ curl -X PUT http://localhost:6333/collections/test_collection4 \
|
||||
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -1419,7 +1416,7 @@ curl -X GET http://localhost:6333/collections/test_collection2/aliases
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.get_collection_aliases(collection_name="{collection_name}")
|
||||
```
|
||||
@@ -1472,7 +1469,7 @@ curl -X GET http://localhost:6333/aliases
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.get_aliases()
|
||||
```
|
||||
@@ -1525,7 +1522,7 @@ curl -X GET http://localhost:6333/collections
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.get_collections()
|
||||
```
|
||||
|
||||
@@ -36,10 +36,9 @@ POST /collections/{collection_name}/points/recommend
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.recommend(
|
||||
collection_name="{collection_name}",
|
||||
@@ -455,12 +454,11 @@ POST /collections/{collection_name}/points/recommend/batch
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
filter = models.Filter(
|
||||
filter_ = models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="city",
|
||||
@@ -473,9 +471,9 @@ filter = models.Filter(
|
||||
|
||||
recommend_queries = [
|
||||
models.RecommendRequest(
|
||||
positive=[100, 231], negative=[718], filter=filter, limit=3
|
||||
positive=[100, 231], negative=[718], filter=filter_, limit=3
|
||||
),
|
||||
models.RecommendRequest(positive=[200, 67], negative=[300], filter=filter, limit=3),
|
||||
models.RecommendRequest(positive=[200, 67], negative=[300], filter=filter_, limit=3),
|
||||
]
|
||||
|
||||
client.recommend_batch(collection_name="{collection_name}", requests=recommend_queries)
|
||||
@@ -703,10 +701,9 @@ POST /collections/{collection_name}/points/discover
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
discover_queries = [
|
||||
models.DiscoverRequest(
|
||||
@@ -906,10 +903,9 @@ POST /collections/{collection_name}/points/discover
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
discover_queries = [
|
||||
models.DiscoverRequest(
|
||||
|
||||
@@ -52,10 +52,9 @@ POST /collections/{collection_name}/points/scroll
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(host="localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.scroll(
|
||||
collection_name="{collection_name}",
|
||||
@@ -729,7 +728,7 @@ Example:
|
||||
```
|
||||
|
||||
```python
|
||||
FieldCondition(
|
||||
models.FieldCondition(
|
||||
key="color",
|
||||
match=models.MatchAny(any=["black", "yellow"]),
|
||||
)
|
||||
@@ -785,7 +784,7 @@ Example:
|
||||
```
|
||||
|
||||
```python
|
||||
FieldCondition(
|
||||
models.FieldCondition(
|
||||
key="color",
|
||||
match=models.MatchExcept(**{"except": ["black", "yellow"]}),
|
||||
)
|
||||
|
||||
@@ -36,7 +36,7 @@ PUT /collections/{collection_name}/index
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(host="localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
@@ -141,10 +141,9 @@ PUT /collections/{collection_name}/index
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(host="localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
@@ -316,16 +315,15 @@ PUT /collections/{collection_name}/index
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(host="localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="name_of_the_field_to_index",
|
||||
field_schema=models.IntegerIndexParams(
|
||||
type="integer",
|
||||
type=models.IntegerIndexType.INTEGER,
|
||||
lookup=False,
|
||||
range=True,
|
||||
),
|
||||
|
||||
@@ -200,10 +200,9 @@ PUT /collections/{collection_name}/points
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(host="localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
@@ -610,6 +609,49 @@ await client.SetPayloadAsync(
|
||||
);
|
||||
```
|
||||
|
||||
_Available as of v1.8.0_
|
||||
|
||||
It is possible to modify only a specific key of the payload by using the `key` parameter.
|
||||
|
||||
For instance, given the following payload JSON object on a point:
|
||||
|
||||
```json
|
||||
{
|
||||
"property1": {
|
||||
"nested_property": "foo",
|
||||
},
|
||||
"property2": {
|
||||
"nested_property": "bar",
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
You can modify the `nested_property` of `property1` with the following request:
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/payload
|
||||
{
|
||||
"payload": {
|
||||
"nested_property": "qux",
|
||||
},
|
||||
"key": "property1",
|
||||
"points": [1]
|
||||
}
|
||||
```
|
||||
|
||||
Resulting in the following payload:
|
||||
|
||||
```json
|
||||
{
|
||||
"property1": {
|
||||
"nested_property": "qux",
|
||||
},
|
||||
"property2": {
|
||||
"nested_property": "bar",
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Overwrite payload
|
||||
|
||||
Fully replace any existing payload with the given one.
|
||||
@@ -725,9 +767,7 @@ POST /collections/{collection_name}/points/payload/clear
|
||||
```python
|
||||
client.clear_payload(
|
||||
collection_name="{collection_name}",
|
||||
points_selector=models.PointIdsList(
|
||||
points=[0, 3, 100],
|
||||
),
|
||||
points_selector=[0, 3, 100],
|
||||
)
|
||||
```
|
||||
|
||||
|
||||
@@ -82,10 +82,9 @@ PUT /collections/{collection_name}/points
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
@@ -1212,9 +1211,7 @@ POST /collections/{collection_name}/points/vectors/delete
|
||||
```python
|
||||
client.delete_vectors(
|
||||
collection_name="{collection_name}",
|
||||
points_selector=models.PointIdsList(
|
||||
points=[0, 3, 100],
|
||||
),
|
||||
points=[0, 3, 100],
|
||||
vectors=["text", "image"],
|
||||
)
|
||||
```
|
||||
@@ -1955,7 +1952,7 @@ POST /collections/{collection_name}/points/batch
|
||||
|
||||
```python
|
||||
client.batch_update_points(
|
||||
collection_name=collection_name,
|
||||
collection_name="{collection_name}",
|
||||
update_operations=[
|
||||
models.UpsertOperation(
|
||||
upsert=models.PointsList(
|
||||
|
||||
@@ -83,10 +83,9 @@ POST /collections/{collection_name}/points/search
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -250,9 +249,8 @@ POST /collections/{collection_name}/points/search
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -358,10 +356,9 @@ POST /collections/{collection_name}/points/search
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -562,9 +559,8 @@ POST /collections/{collection_name}/points/search
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -656,10 +652,9 @@ POST /collections/{collection_name}/points/search
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -808,12 +803,11 @@ POST /collections/{collection_name}/points/search/batch
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
filter = models.Filter(
|
||||
filter_ = models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="city",
|
||||
@@ -825,8 +819,8 @@ filter = models.Filter(
|
||||
)
|
||||
|
||||
search_queries = [
|
||||
models.SearchRequest(vector=[0.2, 0.1, 0.9, 0.7], filter=filter, limit=3),
|
||||
models.SearchRequest(vector=[0.5, 0.3, 0.2, 0.3], filter=filter, limit=3),
|
||||
models.SearchRequest(vector=[0.2, 0.1, 0.9, 0.7], filter=filter_, limit=3),
|
||||
models.SearchRequest(vector=[0.5, 0.3, 0.2, 0.3], filter=filter_, limit=3),
|
||||
]
|
||||
|
||||
client.search_batch(collection_name="{collection_name}", requests=search_queries)
|
||||
@@ -1003,7 +997,7 @@ POST /collections/{collection_name}/points/search
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -1185,7 +1179,7 @@ POST /collections/{collection_name}/points/search/groups
|
||||
client.search_groups(
|
||||
collection_name="{collection_name}",
|
||||
# Same as in the regular search() API
|
||||
query_vector=g,
|
||||
query_vector=[1.1],
|
||||
# Grouping parameters
|
||||
group_by="document_id", # Path of the field to group by
|
||||
limit=4, # Max amount of groups
|
||||
|
||||
@@ -51,7 +51,7 @@ POST /collections/{collection_name}/snapshots
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_snapshot(collection_name="{collection_name}")
|
||||
```
|
||||
@@ -103,7 +103,7 @@ DELETE /collections/{collection_name}/snapshots/{snapshot_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.delete_snapshot(
|
||||
collection_name="{collection_name}", snapshot_name="{snapshot_name}"
|
||||
@@ -155,7 +155,7 @@ GET /collections/{collection_name}/snapshots
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.list_snapshots(collection_name="{collection_name}")
|
||||
```
|
||||
@@ -241,7 +241,7 @@ PUT /collections/{collection_name}/snapshots/recover
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("qdrant-node-2", port=6333)
|
||||
client = QdrantClient(url="http://qdrant-node-2:6333")
|
||||
|
||||
client.recover_snapshot(
|
||||
"{collection_name}",
|
||||
@@ -326,7 +326,7 @@ PUT /collections/{collection_name}/snapshots/recover
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("qdrant-node-2", port=6333)
|
||||
client = QdrantClient(url="http://qdrant-node-2:6333")
|
||||
|
||||
client.recover_snapshot(
|
||||
"{collection_name}",
|
||||
@@ -371,7 +371,7 @@ POST /snapshots
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_full_snapshot()
|
||||
```
|
||||
@@ -421,7 +421,7 @@ DELETE /snapshots/{snapshot_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.delete_full_snapshot(snapshot_name="{snapshot_name}")
|
||||
```
|
||||
|
||||
@@ -56,7 +56,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -168,7 +168,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -293,7 +293,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
|
||||
@@ -17,6 +17,7 @@ be done in the following way:
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
from qdrant_client.models import Batch
|
||||
|
||||
from aleph_alpha_client import (
|
||||
Prompt,
|
||||
@@ -25,7 +26,6 @@ from aleph_alpha_client import (
|
||||
SemanticRepresentation,
|
||||
ImagePrompt
|
||||
)
|
||||
from qdrant_client.http.models import Batch
|
||||
|
||||
aa_token = "<< your_token >>"
|
||||
model = "luminous-base"
|
||||
|
||||
@@ -35,7 +35,7 @@ bedrock_client = session.client(
|
||||
aws_secret_access_key="<YOUR_AWS_SECRET_KEY>",
|
||||
)
|
||||
|
||||
qdrant_client = QdrantClient(location="http://localhost:6333")
|
||||
qdrant_client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
qdrant_client.create_collection(
|
||||
"{collection_name}",
|
||||
|
||||
@@ -18,8 +18,7 @@ The embeddings returned by co.embed API might be used directly in the Qdrant cli
|
||||
```python
|
||||
import cohere
|
||||
import qdrant_client
|
||||
|
||||
from qdrant_client.http.models import Batch
|
||||
from qdrant_client.models import Batch
|
||||
|
||||
cohere_client = cohere.Client("<< your_api_key >>")
|
||||
qdrant_client = qdrant_client.QdrantClient()
|
||||
@@ -55,12 +54,11 @@ documents with the Embed v3 model:
|
||||
```python
|
||||
import cohere
|
||||
import qdrant_client
|
||||
|
||||
from qdrant_client.http.models import Batch
|
||||
from qdrant_client.models import Batch
|
||||
|
||||
cohere_client = cohere.Client("<< your_api_key >>")
|
||||
qdrant_client = qdrant_client.QdrantClient()
|
||||
qdrant_client.upsert(
|
||||
client = qdrant_client.QdrantClient()
|
||||
client.upsert(
|
||||
collection_name="MyCollection",
|
||||
points=Batch(
|
||||
ids=[1],
|
||||
@@ -76,9 +74,9 @@ qdrant_client.upsert(
|
||||
Once the documents are indexed, you can search for the most relevant documents using the Embed v3 model:
|
||||
|
||||
```python
|
||||
qdrant_client.search(
|
||||
client.search(
|
||||
collection_name="MyCollection",
|
||||
query=cohere_client.embed(
|
||||
query_vector=cohere_client.embed(
|
||||
model="embed-english-v3.0", # New Embed v3 model
|
||||
input_type="search_query", # Input type for search queries
|
||||
texts=["The best vector database"],
|
||||
|
||||
@@ -43,11 +43,13 @@ The following example shows how to embed a document with the `models/embedding-0
|
||||
```python
|
||||
import google.generativeai as gemini_client
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http.models import Distance, PointStruct, VectorParams
|
||||
from qdrant_client.models import Distance, PointStruct, VectorParams
|
||||
|
||||
collection_name = "example_collection"
|
||||
|
||||
GEMINI_API_KEY = "YOUR GEMINI API KEY" # add your key here
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
gemini_client.configure(api_key=GEMINI_API_KEY)
|
||||
texts = [
|
||||
"Qdrant is a vector database that is compatible with Gemini.",
|
||||
@@ -83,7 +85,7 @@ points = [
|
||||
### Create Collection
|
||||
|
||||
```python
|
||||
search_client.create_collection(collection_name, vectors_config=
|
||||
client.create_collection(collection_name, vectors_config=
|
||||
VectorParams(
|
||||
size=768,
|
||||
distance=Distance.COSINE,
|
||||
@@ -94,7 +96,7 @@ search_client.create_collection(collection_name, vectors_config=
|
||||
### Add these into the collection
|
||||
|
||||
```python
|
||||
search_client.upsert(collection_name, points)
|
||||
client.upsert(collection_name, points)
|
||||
```
|
||||
|
||||
## Searching for documents with Qdrant
|
||||
@@ -102,7 +104,7 @@ search_client.upsert(collection_name, points)
|
||||
Once the documents are indexed, you can search for the most relevant documents using the same model with the `retrieval_query` task type:
|
||||
|
||||
```python
|
||||
search_client.search(
|
||||
client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=gemini_client.embed_content(
|
||||
model="models/embedding-001",
|
||||
|
||||
@@ -16,8 +16,7 @@ To call their endpoint, all you need is an API key obtainable [here](https://jin
|
||||
import qdrant_client
|
||||
import requests
|
||||
|
||||
from qdrant_client.http.models import Distance, VectorParams
|
||||
from qdrant_client.http.models import Batch
|
||||
from qdrant_client.models import Distance, VectorParams, Batch
|
||||
|
||||
# Provide Jina API key and choose one of the available models.
|
||||
# You can get a free trial key here: https://jina.ai/embeddings/
|
||||
@@ -43,8 +42,8 @@ embeddings = [d["embedding"] for d in response.json()["data"]]
|
||||
|
||||
|
||||
# Index the embeddings into Qdrant
|
||||
qdrant_client = qdrant_client.QdrantClient(":memory:")
|
||||
qdrant_client.create_collection(
|
||||
client = qdrant_client.QdrantClient(":memory:")
|
||||
client.create_collection(
|
||||
collection_name="MyCollection",
|
||||
vectors_config=VectorParams(size=EMBEDDING_SIZE, distance=Distance.DOT),
|
||||
)
|
||||
|
||||
@@ -22,11 +22,12 @@ And then we set this up:
|
||||
```python
|
||||
from mistralai.client import MistralClient
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http.models import PointStruct, VectorParams, Distance
|
||||
from qdrant_client.models import PointStruct, VectorParams, Distance
|
||||
|
||||
collection_name = "example_collection"
|
||||
|
||||
MISTRAL_API_KEY = "your_mistral_api_key"
|
||||
search_client = QdrantClient(":memory:")
|
||||
client = QdrantClient(":memory:")
|
||||
mistral_client = MistralClient(api_key=MISTRAL_API_KEY)
|
||||
texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
@@ -65,13 +66,12 @@ points = [
|
||||
## Create a collection and Insert the documents
|
||||
|
||||
```python
|
||||
search_client.create_collection(collection_name, vectors_config=
|
||||
VectorParams(
|
||||
client.create_collection(collection_name, vectors_config=VectorParams(
|
||||
size=1024,
|
||||
distance=Distance.COSINE,
|
||||
)
|
||||
)
|
||||
search_client.upsert(collection_name, points)
|
||||
client.upsert(collection_name, points)
|
||||
```
|
||||
|
||||
## Searching for documents with Qdrant
|
||||
@@ -79,7 +79,7 @@ search_client.upsert(collection_name, points)
|
||||
Once the documents are indexed, you can search for the most relevant documents using the same model with the `retrieval_query` task type:
|
||||
|
||||
```python
|
||||
search_client.search(
|
||||
client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=mistral_client.embeddings(
|
||||
model="mistral-embed", input=["What is the best to use for vector search scaling?"]
|
||||
|
||||
@@ -30,8 +30,8 @@ output = embed.text(
|
||||
task_type="search_document",
|
||||
)
|
||||
|
||||
qdrant_client = QdrantClient()
|
||||
qdrant_client.upsert(
|
||||
client = QdrantClient()
|
||||
client.upsert(
|
||||
collection_name="my-collection",
|
||||
points=models.Batch(
|
||||
ids=[1],
|
||||
@@ -44,14 +44,14 @@ qdrant_client.upsert(
|
||||
|
||||
```python
|
||||
from fastembed import TextEmbedding
|
||||
from qdrant_client import QdrantClient, models
|
||||
from client import QdrantClient, models
|
||||
|
||||
model = TextEmbedding("nomic-ai/nomic-embed-text-v1")
|
||||
|
||||
output = model.embed(["Qdrant is the best vector database!"])
|
||||
|
||||
qdrant_client = QdrantClient()
|
||||
qdrant_client.upsert(
|
||||
client = QdrantClient()
|
||||
client.upsert(
|
||||
collection_name="my-collection",
|
||||
points=models.Batch(
|
||||
ids=[1],
|
||||
@@ -71,7 +71,7 @@ output = embed.text(
|
||||
task_type="search_query",
|
||||
)
|
||||
|
||||
qdrant_client.search(
|
||||
client.search(
|
||||
collection_name="my-collection",
|
||||
query_vector=output["embeddings"][0],
|
||||
)
|
||||
@@ -82,7 +82,7 @@ qdrant_client.search(
|
||||
```python
|
||||
output = next(model.embed("What is the best vector database?"))
|
||||
|
||||
qdrant_client.search(
|
||||
client.search(
|
||||
collection_name="my-collection",
|
||||
query_vector=output.tolist(),
|
||||
)
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
---
|
||||
title: Nvidia
|
||||
weight: 1200
|
||||
---
|
||||
|
||||
# Nvidia
|
||||
|
||||
Qdrant supports working with [Nvidia embeddings](https://build.nvidia.com/explore/retrieval).
|
||||
|
||||
You can generate an API key to authenticate the requests from the [Nvidia Playground](<https://build.nvidia.com/nvidia/embed-qa-4>).
|
||||
|
||||
### Setting up the Qdrant client and Nvidia session
|
||||
|
||||
```python
|
||||
import requests
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
NVIDIA_BASE_URL = "https://ai.api.nvidia.com/v1/retrieval/nvidia/embeddings"
|
||||
|
||||
NVIDIA_API_KEY = "<YOUR_API_KEY>"
|
||||
|
||||
nvidia_session = requests.Session()
|
||||
|
||||
client = QdrantClient(":memory:")
|
||||
|
||||
headers = {
|
||||
"Authorization": f"Bearer {NVIDIA_API_KEY}",
|
||||
"Accept": "application/json",
|
||||
}
|
||||
|
||||
texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from '@qdrant/js-client-rest';
|
||||
|
||||
const NVIDIA_BASE_URL = "https://ai.api.nvidia.com/v1/retrieval/nvidia/embeddings"
|
||||
const NVIDIA_API_KEY = "<YOUR_API_KEY>"
|
||||
|
||||
const client = new QdrantClient({ url: 'http://localhost:6333' });
|
||||
|
||||
const headers = {
|
||||
"Authorization": "Bearer " + NVIDIA_API_KEY,
|
||||
"Accept": "application/json",
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
const texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
The following example shows how to embed documents with the `embed-qa-4` model that generates sentence embeddings of size 1024.
|
||||
|
||||
### Embedding documents
|
||||
|
||||
```python
|
||||
payload = {
|
||||
"input": texts,
|
||||
"input_type": "passage",
|
||||
"model": "NV-Embed-QA",
|
||||
}
|
||||
|
||||
response_body = nvidia_session.post(
|
||||
NVIDIA_BASE_URL, headers=headers, json=payload
|
||||
).json()
|
||||
```
|
||||
|
||||
```typescript
|
||||
let body = {
|
||||
"input": texts,
|
||||
"input_type": "passage",
|
||||
"model": "NV-Embed-QA"
|
||||
}
|
||||
|
||||
let response = await fetch(NVIDIA_BASE_URL, {
|
||||
method: "POST",
|
||||
body: JSON.stringify(body),
|
||||
headers
|
||||
});
|
||||
|
||||
let response_body = await response.json()
|
||||
```
|
||||
|
||||
### Converting the model outputs to Qdrant points
|
||||
|
||||
```python
|
||||
from qdrant_client.models import PointStruct
|
||||
|
||||
points = [
|
||||
PointStruct(
|
||||
id=idx,
|
||||
vector=data["embedding"],
|
||||
payload={"text": text},
|
||||
)
|
||||
for idx, (data, text) in enumerate(zip(response_body["data"], texts))
|
||||
]
|
||||
```
|
||||
|
||||
```typescript
|
||||
let points = response_body.data.map((data, i) => {
|
||||
return {
|
||||
id: i,
|
||||
vector: data.embedding,
|
||||
payload: {
|
||||
text: texts[i]
|
||||
}
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
### Creating a collection to insert the documents
|
||||
|
||||
```python
|
||||
from qdrant_client.models import VectorParams, Distance
|
||||
|
||||
collection_name = "example_collection"
|
||||
|
||||
client.create_collection(
|
||||
collection_name,
|
||||
vectors_config=VectorParams(
|
||||
size=1024,
|
||||
distance=Distance.COSINE,
|
||||
),
|
||||
)
|
||||
client.upsert(collection_name, points)
|
||||
```
|
||||
|
||||
```typescript
|
||||
const COLLECTION_NAME = "example_collection"
|
||||
|
||||
await client.createCollection(COLLECTION_NAME, {
|
||||
vectors: {
|
||||
size: 1024,
|
||||
distance: 'Cosine',
|
||||
}
|
||||
});
|
||||
|
||||
await client.upsert(COLLECTION_NAME, {
|
||||
wait: true,
|
||||
points
|
||||
})
|
||||
```
|
||||
|
||||
## Searching for documents with Qdrant
|
||||
|
||||
Once the documents are added, you can search for the most relevant documents.
|
||||
|
||||
```python
|
||||
payload = {
|
||||
"input": "What is the best to use for vector search scaling?",
|
||||
"input_type": "query",
|
||||
"model": "NV-Embed-QA",
|
||||
}
|
||||
|
||||
response_body = nvidia_session.post(
|
||||
NVIDIA_BASE_URL, headers=headers, json=payload
|
||||
).json()
|
||||
|
||||
client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=response_body["data"][0]["embedding"],
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
body = {
|
||||
"input": "What is the best to use for vector search scaling?",
|
||||
"input_type": "query",
|
||||
"model": "NV-Embed-QA",
|
||||
}
|
||||
|
||||
response = await fetch(NVIDIA_BASE_URL, {
|
||||
method: "POST",
|
||||
body: JSON.stringify(body),
|
||||
headers
|
||||
});
|
||||
|
||||
response_body = await response.json()
|
||||
|
||||
await client.search(COLLECTION_NAME, {
|
||||
vector: response_body.data[0].embedding,
|
||||
});
|
||||
```
|
||||
@@ -24,7 +24,7 @@ openai_client = openai.Client(
|
||||
api_key="<YOUR_API_KEY>"
|
||||
)
|
||||
|
||||
qdrant_client = qdrant_client.QdrantClient(":memory:")
|
||||
client = qdrant_client.QdrantClient(":memory:")
|
||||
|
||||
texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
@@ -39,13 +39,13 @@ The following example shows how to embed a document with the `text-embedding-3-s
|
||||
```python
|
||||
embedding_model = "text-embedding-3-small"
|
||||
|
||||
result = openai_client.embeddings.create(input= texts, model=embedding_model)
|
||||
result = openai_client.embeddings.create(input=texts, model=embedding_model)
|
||||
```
|
||||
|
||||
### Converting the model outputs to Qdrant points
|
||||
|
||||
```python
|
||||
from qdrant_client.http.models import PointStruct
|
||||
from qdrant_client.models import PointStruct
|
||||
|
||||
points = [
|
||||
PointStruct(
|
||||
@@ -60,18 +60,18 @@ points = [
|
||||
### Creating a collection to insert the documents
|
||||
|
||||
```python
|
||||
from qdrant_client.http.models import VectorParams, Distance
|
||||
from qdrant_client.models import VectorParams, Distance
|
||||
|
||||
collection_name = "example_collection"
|
||||
|
||||
qdrant_client.create_collection(
|
||||
client.create_collection(
|
||||
collection_name,
|
||||
vectors_config=VectorParams(
|
||||
size=1536,
|
||||
distance=Distance.COSINE,
|
||||
),
|
||||
)
|
||||
qdrant_client.upsert(collection_name, points)
|
||||
client.upsert(collection_name, points)
|
||||
```
|
||||
|
||||
## Searching for documents with Qdrant
|
||||
@@ -79,7 +79,7 @@ qdrant_client.upsert(collection_name, points)
|
||||
Once the documents are indexed, you can search for the most relevant documents using the same model.
|
||||
|
||||
```python
|
||||
qdrant_client.search(
|
||||
client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=openai_client.embeddings.create(
|
||||
input=["What is the best to use for vector search scaling?"],
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
---
|
||||
title: Voyage AI
|
||||
weight: 1300
|
||||
---
|
||||
|
||||
# Voyage AI
|
||||
|
||||
Qdrant supports working with [Voyage AI](https://voyageai.com/) embeddings. The supported models' list can be found [here](https://docs.voyageai.com/docs/embeddings).
|
||||
|
||||
You can generate an API key from the [Voyage AI dashboard](<https://dash.voyageai.com/>) to authenticate the requests.
|
||||
|
||||
### Setting up the Qdrant and Voyage clients
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
import voyageai
|
||||
|
||||
VOYAGE_API_KEY = "<YOUR_VOYAGEAI_API_KEY>"
|
||||
|
||||
qclient = QdrantClient(":memory:")
|
||||
vclient = voyageai.Client(api_key=VOYAGE_API_KEY)
|
||||
|
||||
texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
```typescript
|
||||
import {QdrantClient} from '@qdrant/js-client-rest';
|
||||
|
||||
const VOYAGEAI_BASE_URL = "https://api.voyageai.com/v1/embeddings"
|
||||
const VOYAGEAI_API_KEY = "<YOUR_VOYAGEAI_API_KEY>"
|
||||
|
||||
const client = new QdrantClient({ url: 'http://localhost:6333' });
|
||||
|
||||
const headers = {
|
||||
"Authorization": "Bearer " + VOYAGEAI_API_KEY,
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
const texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
The following example shows how to embed documents with the [`voyage-large-2`](https://docs.voyageai.com/docs/embeddings#model-choices) model that generates sentence embeddings of size 1536.
|
||||
|
||||
### Embedding documents
|
||||
|
||||
```python
|
||||
response = vclient.embed(texts, model="voyage-large-2", input_type="document")
|
||||
```
|
||||
|
||||
```typescript
|
||||
let body = {
|
||||
"input": texts,
|
||||
"model": "voyage-large-2",
|
||||
"input_type": "document",
|
||||
}
|
||||
|
||||
let response = await fetch(VOYAGEAI_BASE_URL, {
|
||||
method: "POST",
|
||||
body: JSON.stringify(body),
|
||||
headers
|
||||
});
|
||||
|
||||
let response_body = await response.json();
|
||||
```
|
||||
|
||||
### Converting the model outputs to Qdrant points
|
||||
|
||||
```python
|
||||
from qdrant_client.models import PointStruct
|
||||
|
||||
points = [
|
||||
PointStruct(
|
||||
id=idx,
|
||||
vector=embedding,
|
||||
payload={"text": text},
|
||||
)
|
||||
for idx, (embedding, text) in enumerate(zip(response.embeddings, texts))
|
||||
]
|
||||
```
|
||||
|
||||
```typescript
|
||||
let points = response_body.data.map((data, i) => {
|
||||
return {
|
||||
id: i,
|
||||
vector: data.embedding,
|
||||
payload: {
|
||||
text: texts[i]
|
||||
}
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
### Creating a collection to insert the documents
|
||||
|
||||
```python
|
||||
from qdrant_client.models import VectorParams, Distance
|
||||
|
||||
COLLECTION_NAME = "example_collection"
|
||||
|
||||
qclient.create_collection(
|
||||
COLLECTION_NAME,
|
||||
vectors_config=VectorParams(
|
||||
size=1536,
|
||||
distance=Distance.COSINE,
|
||||
),
|
||||
)
|
||||
qclient.upsert(COLLECTION_NAME, points)
|
||||
```
|
||||
|
||||
```typescript
|
||||
const COLLECTION_NAME = "example_collection"
|
||||
|
||||
await client.createCollection(COLLECTION_NAME, {
|
||||
vectors: {
|
||||
size: 1536,
|
||||
distance: 'Cosine',
|
||||
}
|
||||
});
|
||||
|
||||
await client.upsert(COLLECTION_NAME, {
|
||||
wait: true,
|
||||
points
|
||||
});
|
||||
```
|
||||
|
||||
### Searching for documents with Qdrant
|
||||
|
||||
Once the documents are added, you can search for the most relevant documents.
|
||||
|
||||
```python
|
||||
response = vclient.embed(
|
||||
["What is the best to use for vector search scaling?"],
|
||||
model="voyage-large-2",
|
||||
input_type="query",
|
||||
)
|
||||
|
||||
qclient.search(
|
||||
collection_name=COLLECTION_NAME,
|
||||
query_vector=response.embeddings[0],
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
body = {
|
||||
"input": ["What is the best to use for vector search scaling?"],
|
||||
"model": "voyage-large-2",
|
||||
"input_type": "query",
|
||||
};
|
||||
|
||||
response = await fetch(VOYAGEAI_BASE_URL, {
|
||||
method: "POST",
|
||||
body: JSON.stringify(body),
|
||||
headers
|
||||
});
|
||||
|
||||
response_body = await response.json();
|
||||
|
||||
await client.search(COLLECTION_NAME, {
|
||||
vector: response_body.data[0].embedding,
|
||||
});
|
||||
```
|
||||
@@ -69,7 +69,7 @@ ids = [32, 21, "b626f6a9-b14d-4af9-b7c3-43d8deb719a6"]
|
||||
payload = [{"meta": "data"}, {"meta": "data_2"}, {"meta": "data_3", "extra": "data"}]
|
||||
|
||||
QdrantIngestOperator(
|
||||
conn_id="qdrant_connection"
|
||||
conn_id="qdrant_connection",
|
||||
task_id="qdrant_ingest",
|
||||
collection_name="<COLLECTION_NAME>",
|
||||
vectors=vectors,
|
||||
|
||||
@@ -68,7 +68,7 @@ assistant = RetrieveAssistantAgent(
|
||||
# `chunk_token_size` is the chunk token size for the retrieve chat.
|
||||
# We use an in-memory QdrantClient instance here. Not recommended for production.
|
||||
|
||||
ragproxyagent = QdrantRetrieveUserProxyAgent(
|
||||
rag_proxy_agent = QdrantRetrieveUserProxyAgent(
|
||||
name="qdrantagent",
|
||||
human_input_mode="NEVER",
|
||||
max_consecutive_auto_reply=10,
|
||||
@@ -95,7 +95,7 @@ assistant.reset()
|
||||
|
||||
# The query used below is for demonstration. It should usually be related to the docs made available to the agent
|
||||
code_problem = "How can I use FLAML to perform a classification task?"
|
||||
ragproxyagent.initiate_chat(assistant, problem=code_problem)
|
||||
rag_proxy_agent.initiate_chat(assistant, problem=code_problem)
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
title: Pinecone Canopy
|
||||
weight: 2500
|
||||
---
|
||||
|
||||
# Pinecone Canopy
|
||||
|
||||
[Canopy](https://github.com/pinecone-io/canopy) is an open-source framework and context engine to build chat assistants at scale.
|
||||
|
||||
Qdrant is supported as a knowledge base within Canopy for context retrieval and augmented generation.
|
||||
|
||||
## Usage
|
||||
|
||||
Install the SDK with the Qdrant extra as described in the [Canopy README](https://github.com/pinecone-io/canopy?tab=readme-ov-file#extras).
|
||||
|
||||
```bash
|
||||
pip install canopy-sdk[qdrant]
|
||||
```
|
||||
|
||||
### Creating a knowledge base
|
||||
|
||||
```python
|
||||
from canopy.knowledge_base import QdrantKnowledgeBase
|
||||
|
||||
kb = QdrantKnowledgeBase(collection_name="<YOUR_COLLECTION_NAME>")
|
||||
```
|
||||
|
||||
<aside role="status">The constructor accepts additional <a href="https://github.com/qdrant/qdrant-client/blob/eda201a1dbf1bbc67415f8437a5619f6f83e8ac6/qdrant_client/qdrant_client.py#L36-L61">options</a> to customize your connection to Qdrant.</aside>
|
||||
|
||||
To create a new Qdrant collection and connect it to the knowledge base, use the `create_canopy_collection` method:
|
||||
|
||||
```python
|
||||
kb.create_canopy_collection()
|
||||
```
|
||||
|
||||
You can always verify the connection to the collection with the `verify_index_connection` method:
|
||||
|
||||
```python
|
||||
kb.verify_index_connection()
|
||||
```
|
||||
|
||||
Learn more about customizing the knowledge base and its inner components [in the Canopy library](https://github.com/pinecone-io/canopy/blob/main/docs/library.md#understanding-knowledgebase-workings).
|
||||
|
||||
### Adding data to the knowledge base
|
||||
|
||||
To insert data into the knowledge base, you can create a list of documents and use the `upsert` method:
|
||||
|
||||
```python
|
||||
from canopy.models.data_models import Document
|
||||
|
||||
documents = [
|
||||
Document(
|
||||
id="1",
|
||||
text="U2 are an Irish rock band from Dublin, formed in 1976.",
|
||||
source="https://en.wikipedia.org/wiki/U2",
|
||||
),
|
||||
Document(
|
||||
id="2",
|
||||
text="Arctic Monkeys are an English rock band formed in Sheffield in 2002.",
|
||||
source="https://en.wikipedia.org/wiki/Arctic_Monkeys",
|
||||
metadata={"my-key": "my-value"},
|
||||
),
|
||||
]
|
||||
|
||||
kb.upsert(documents)
|
||||
```
|
||||
|
||||
### Querying the knowledge base
|
||||
|
||||
You can query the knowledge base with the `query` method to find the most similar documents to a given text:
|
||||
|
||||
```python
|
||||
from canopy.models.data_models import Query
|
||||
|
||||
kb.query(
|
||||
[
|
||||
Query(text="Arctic Monkeys music genre"),
|
||||
Query(
|
||||
text="U2 music genre",
|
||||
top_k=10,
|
||||
metadata_filter={"key": "my-key", "match": {"value": "my-value"}},
|
||||
),
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Introduction to Canopy](https://www.pinecone.io/blog/canopy-rag-framework/)
|
||||
- [Canopy library reference](https://github.com/pinecone-io/canopy/blob/main/docs/library.md)
|
||||
@@ -27,7 +27,7 @@ Scalar Quantization, you'd make that in the following way:
|
||||
|
||||
```python
|
||||
from qdrant_haystack.document_stores import QdrantDocumentStore
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import models
|
||||
|
||||
document_store = QdrantDocumentStore(
|
||||
":memory:",
|
||||
|
||||
@@ -68,7 +68,8 @@ client is destroyed - usually at the end of your script/notebook.
|
||||
|
||||
```python
|
||||
qdrant = Qdrant.from_documents(
|
||||
docs, embeddings,
|
||||
docs,
|
||||
embeddings,
|
||||
location=":memory:", # Local mode with in-memory storage only
|
||||
collection_name="my_documents",
|
||||
)
|
||||
@@ -80,7 +81,8 @@ Local mode, without using the Qdrant server, may also store your vectors on disk
|
||||
|
||||
```python
|
||||
qdrant = Qdrant.from_documents(
|
||||
docs, embeddings,
|
||||
docs,
|
||||
embeddings,
|
||||
path="/tmp/local_qdrant",
|
||||
collection_name="my_documents",
|
||||
)
|
||||
|
||||
@@ -36,5 +36,5 @@ index = VectorStoreIndex.from_vector_store(vector_store=vector_store)
|
||||
|
||||
```
|
||||
|
||||
The library [comes with a notebook](https://github.com/run-llama/llama_index/blob/main/docs/examples/vector_stores/QdrantIndexDemo.ipynb)
|
||||
The library [comes with a notebook](https://colab.research.google.com/github/run-llama/llama_index/blob/main/docs/docs/examples/vector_stores/QdrantIndexDemo.ipynb)
|
||||
that shows an end-to-end example of how to use Qdrant within LlamaIndex.
|
||||
|
||||
@@ -0,0 +1,95 @@
|
||||
---
|
||||
title: Pandas-AI
|
||||
weight: 2900
|
||||
---
|
||||
|
||||
# Pandas-AI
|
||||
|
||||
Pandas-AI is a Python library that uses a generative AI model to interpret natural language queries and translate them into Python code to interact with pandas data frames and return the final results to the user.
|
||||
|
||||
## Installation
|
||||
|
||||
```console
|
||||
pip install pandasai[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
You can begin a conversation by instantiating an `Agent` instance based on your Pandas data frame. The default Pandas-AI LLM requires an [API key](https://pandabi.ai).
|
||||
|
||||
You can find the list of all supported LLMs [here](https://docs.pandas-ai.com/en/latest/LLMs/llms/)
|
||||
|
||||
```python
|
||||
import os
|
||||
import pandas as pd
|
||||
from pandasai import Agent
|
||||
|
||||
# Sample DataFrame
|
||||
sales_by_country = pd.DataFrame(
|
||||
{
|
||||
"country": [
|
||||
"United States",
|
||||
"United Kingdom",
|
||||
"France",
|
||||
"Germany",
|
||||
"Italy",
|
||||
"Spain",
|
||||
"Canada",
|
||||
"Australia",
|
||||
"Japan",
|
||||
"China",
|
||||
],
|
||||
"sales": [5000, 3200, 2900, 4100, 2300, 2100, 2500, 2600, 4500, 7000],
|
||||
}
|
||||
)
|
||||
|
||||
os.environ["PANDASAI_API_KEY"] = "YOUR_API_KEY"
|
||||
|
||||
agent = Agent(sales_by_country)
|
||||
agent.chat("Which are the top 5 countries by sales?")
|
||||
# OUTPUT: China, United States, Japan, Germany, Australia
|
||||
```
|
||||
|
||||
## Qdrant support
|
||||
|
||||
You can train Pandas-AI to understand your data better and improve the quality of the results.
|
||||
|
||||
Qdrant can be configured as a vector store to ingest training data and retrieve semantically relevant content.
|
||||
|
||||
```python
|
||||
from pandasai.ee.vectorstores.qdrant import Qdrant
|
||||
|
||||
qdrant = Qdrant(
|
||||
collection_name="<SOME_COLLECTION>",
|
||||
embedding_model="sentence-transformers/all-MiniLM-L6-v2",
|
||||
url="http://localhost:6333",
|
||||
grpc_port=6334,
|
||||
prefer_grpc=True
|
||||
)
|
||||
|
||||
agent = Agent(df, vector_store=qdrant)
|
||||
|
||||
# Train with custom information
|
||||
agent.train(docs="The fiscal year starts in April")
|
||||
|
||||
# Train the q/a pairs of code snippets
|
||||
query = "What are the total sales for the current fiscal year?"
|
||||
response = """
|
||||
import pandas as pd
|
||||
|
||||
df = dfs[0]
|
||||
|
||||
# Calculate the total sales for the current fiscal year
|
||||
total_sales = df[df['date'] >= pd.to_datetime('today').replace(month=4, day=1)]['sales'].sum()
|
||||
result = { "type": "number", "value": total_sales }
|
||||
"""
|
||||
agent.train(queries=[query], codes=[response])
|
||||
|
||||
# # The model will use the information provided in the training to generate a response
|
||||
|
||||
```
|
||||
|
||||
## Further reading
|
||||
|
||||
- [Getting Started with Pandas-AI](https://pandasai-docs.readthedocs.io/en/latest/getting-started/)
|
||||
- [Pandas-AI Reference](https://pandasai-docs.readthedocs.io/en/latest/)
|
||||
@@ -0,0 +1,105 @@
|
||||
---
|
||||
title: Semantic-Router
|
||||
weight: 2700
|
||||
---
|
||||
|
||||
# Semantic-Router
|
||||
|
||||
[Semantic-Router](https://www.aurelio.ai/semantic-router/) is a library to build decision-making layers for your LLMs and agents. It uses vector embeddings to make tool-use decisions rather than LLM generations, routing our requests using semantic meaning.
|
||||
|
||||
Qdrant is available as a supported index in Semantic-Router for you to ingest route data and perform retrievals.
|
||||
|
||||
## Installation
|
||||
|
||||
To use Semantic-Router with Qdrant, install the `qdrant` extra:
|
||||
|
||||
```console
|
||||
pip install semantic-router[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Set up `QdrantIndex` with the appropriate configurations:
|
||||
|
||||
```python
|
||||
from semantic_router.index import QdrantIndex
|
||||
|
||||
qdrant_index = QdrantIndex(
|
||||
url="https://xyz-example.eu-central.aws.cloud.qdrant.io", api_key="<your-api-key>"
|
||||
)
|
||||
```
|
||||
|
||||
Once the Qdrant index is set up with the appropriate configurations, we can pass it to the `RouteLayer`.
|
||||
|
||||
```python
|
||||
from semantic_router.layer import RouteLayer
|
||||
|
||||
RouteLayer(encoder=some_encoder, routes=some_routes, index=qdrant_index)
|
||||
```
|
||||
|
||||
## Complete Example
|
||||
|
||||
<details>
|
||||
|
||||
<summary><b>Click to expand</b></summary>
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
from semantic_router import Route
|
||||
from semantic_router.encoders import OpenAIEncoder
|
||||
from semantic_router.index import QdrantIndex
|
||||
from semantic_router.layer import RouteLayer
|
||||
|
||||
# we could use this as a guide for our chatbot to avoid political conversations
|
||||
politics = Route(
|
||||
name="politics value",
|
||||
utterances=[
|
||||
"isn't politics the best thing ever",
|
||||
"why don't you tell me about your political opinions",
|
||||
"don't you just love the president",
|
||||
"they're going to destroy this country!",
|
||||
"they will save the country!",
|
||||
],
|
||||
)
|
||||
|
||||
# this could be used as an indicator to our chatbot to switch to a more
|
||||
# conversational prompt
|
||||
chitchat = Route(
|
||||
name="chitchat",
|
||||
utterances=[
|
||||
"how's the weather today?",
|
||||
"how are things going?",
|
||||
"lovely weather today",
|
||||
"the weather is horrendous",
|
||||
"let's go to the chippy",
|
||||
],
|
||||
)
|
||||
|
||||
# we place both of our decisions together into single list
|
||||
routes = [politics, chitchat]
|
||||
|
||||
os.environ["OPENAI_API_KEY"] = "<YOUR_API_KEY>"
|
||||
encoder = OpenAIEncoder()
|
||||
|
||||
rl = RouteLayer(
|
||||
encoder=encoder,
|
||||
routes=routes,
|
||||
index=QdrantIndex(location=":memory:"),
|
||||
)
|
||||
|
||||
print(rl("What have you been upto?").name)
|
||||
```
|
||||
|
||||
This returns:
|
||||
|
||||
```console
|
||||
[Out]: 'chitchat'
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- Semantic-Router [Documentation](https://github.com/aurelio-labs/semantic-router/tree/main/docs)
|
||||
- Semantic-Router [Video Course](https://www.aurelio.ai/course/semantic-router)
|
||||
@@ -244,5 +244,6 @@ Qdrant supports all the Spark data types, and the appropriate data types are map
|
||||
| `sparse_vector_index_fields` | Comma-separated names of columns holding the sparse vector indices. | `ArrayType(IntegerType)` | ❌ |
|
||||
| `sparse_vector_value_fields` | Comma-separated names of columns holding the sparse vector values. | `ArrayType(FloatType)` | ❌ |
|
||||
| `sparse_vector_names` | Comma-separated names of the sparse vectors in the collection. | - | ❌ |
|
||||
| `shard_key_selector` | Comma-separated names of custom shard keys to use during upsert. | - | ❌ |
|
||||
|
||||
For more information, be sure to check out the [Qdrant-Spark GitHub repository](https://github.com/qdrant/qdrant-spark). The Apache Spark guide is available [here](https://spark.apache.org/docs/latest/quick-start.html). Happy data processing!
|
||||
|
||||
@@ -26,7 +26,6 @@ import (
|
||||
qdrantContainer, err := qdrant.RunContainer(ctx, testcontainers.WithImage("qdrant/qdrant"))
|
||||
```
|
||||
|
||||
<!--
|
||||
```typescript
|
||||
import { QdrantContainer } from "@testcontainers/qdrant";
|
||||
|
||||
@@ -38,7 +37,6 @@ from testcontainers.qdrant import QdrantContainer
|
||||
|
||||
qdrant_container = QdrantContainer("qdrant/qdrant").start()
|
||||
```
|
||||
-->
|
||||
|
||||
Testcontainers modules provide options/methods to configure ENVs, volumes, and virtually everything you can configure in a Docker container.
|
||||
|
||||
|
||||
@@ -38,7 +38,7 @@ unstructured-ingest \
|
||||
--verbose \
|
||||
qdrant \
|
||||
--collection-name "test" \
|
||||
--location "http://localhost:6333" \
|
||||
--url "http://localhost:6333" \
|
||||
--batch-size 80
|
||||
```
|
||||
|
||||
@@ -66,7 +66,7 @@ from unstructured.ingest.runner.writers.qdrant import QdrantWriter
|
||||
def get_writer() -> Writer:
|
||||
return QdrantWriter(
|
||||
connector_config=SimpleQdrantConfig(
|
||||
location="http://localhost:6333",
|
||||
url="http://localhost:6333",
|
||||
collection_name="test",
|
||||
),
|
||||
write_config=QdrantWriteConfig(batch_size=80),
|
||||
|
||||
@@ -186,10 +186,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -337,10 +336,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -445,10 +443,9 @@ PUT /collections/{collection_name}/points
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
@@ -662,10 +659,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -913,10 +909,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -1209,7 +1204,7 @@ client.upsert(
|
||||
[0.1, 0.1, 0.9],
|
||||
],
|
||||
),
|
||||
ordering="strong",
|
||||
ordering=models.WriteOrdering.STRONG,
|
||||
)
|
||||
```
|
||||
|
||||
|
||||
@@ -221,7 +221,7 @@ POST /collections/{collection_name}/points/search
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -342,7 +342,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
|
||||
@@ -42,10 +42,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -196,10 +195,9 @@ POST /collections/{collection_name}/points/search
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -316,7 +314,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -467,10 +465,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -620,7 +617,7 @@ POST /collections/{collection_name}/points/search
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -734,7 +731,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -852,7 +849,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
|
||||
@@ -167,10 +167,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -335,10 +334,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -475,10 +473,9 @@ PUT /collections/{collection_name}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -630,10 +627,9 @@ POST /collections/{collection_name}/points/search
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http import models
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -784,7 +780,7 @@ POST /collections/{collection_name}/points/search
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -923,7 +919,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -1077,7 +1073,7 @@ POST /collections/{collection_name}/points/search
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search(
|
||||
collection_name="{collection_name}",
|
||||
@@ -1200,7 +1196,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
|
||||
@@ -63,8 +63,7 @@ curl \
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://localhost",
|
||||
port=6333,
|
||||
url="https://localhost:6333",
|
||||
api_key="your_secret_api_key_here",
|
||||
)
|
||||
```
|
||||
@@ -186,8 +185,7 @@ curl -X GET https://localhost:6333
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://localhost",
|
||||
port=6333,
|
||||
url="https://localhost:6333",
|
||||
)
|
||||
```
|
||||
|
||||
|
||||
@@ -39,7 +39,7 @@ Qdrant is now accessible:
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
```
|
||||
|
||||
```typescript
|
||||
@@ -78,7 +78,7 @@ var client = new QdrantClient("localhost", 6334);
|
||||
You will be storing all of your vector data in a Qdrant collection. Let's call it `test_collection`. This collection will be using a dot product distance metric to compare vectors.
|
||||
|
||||
```python
|
||||
from qdrant_client.http.models import Distance, VectorParams
|
||||
from qdrant_client.models import Distance, VectorParams
|
||||
|
||||
client.create_collection(
|
||||
collection_name="test_collection",
|
||||
@@ -135,7 +135,7 @@ await client.CreateCollectionAsync(
|
||||
Let's now add a few vectors with a payload. Payloads are other data you want to associate with the vector:
|
||||
|
||||
```python
|
||||
from qdrant_client.http.models import PointStruct
|
||||
from qdrant_client.models import PointStruct
|
||||
|
||||
operation_info = client.upsert(
|
||||
collection_name="test_collection",
|
||||
@@ -524,7 +524,7 @@ See [payload and vector in the result](../concepts/search/#payload-and-vector-in
|
||||
We can narrow down the results further by filtering by payload. Let's find the closest results that include "London".
|
||||
|
||||
```python
|
||||
from qdrant_client.http.models import Filter, FieldCondition, MatchValue
|
||||
from qdrant_client.models import Filter, FieldCondition, MatchValue
|
||||
|
||||
search_result = client.search(
|
||||
collection_name="test_collection",
|
||||
|
||||
@@ -27,5 +27,6 @@ These tutorials demonstrate different ways you can build vector search into your
|
||||
| [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina |
|
||||
| [Implement custom connector for Cohere RAG](../tutorials/cohere-rag-connector/) | Bring data stored in Qdrant to Cohere RAG | Qdrant, Cohere, FastAPI |
|
||||
| [Private Chatbot for Interactive Learning](../tutorials/student-rag-haystack-red-hat-openshift-hc/) | Speak to the course materials as if you were speaking to the teacher | Qdrant, Red Hat OpenShift, Haystack, Hugging Face, Mistral, Hybrid Cloud |
|
||||
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
||||
|
||||
@@ -70,7 +70,7 @@ from aleph_alpha_client import (
|
||||
from glob import glob
|
||||
|
||||
ids, vectors, payloads = [], [], []
|
||||
async with AsyncClient(token=aa_token) as client:
|
||||
async with AsyncClient(token=aa_token) as aa_client:
|
||||
for i, image_path in enumerate(glob("./val2017/*.jpg")):
|
||||
# Convert the JPEG file into the embedding by calling
|
||||
# Aleph Alpha API
|
||||
@@ -82,7 +82,7 @@ async with AsyncClient(token=aa_token) as client:
|
||||
"compress_to_size": 128,
|
||||
}
|
||||
query_request = SemanticEmbeddingRequest(**query_params)
|
||||
query_response = await client.semantic_embed(request=query_request, model=model)
|
||||
query_response = await aa_client.semantic_embed(request=query_request, model=model)
|
||||
|
||||
# Finally store the id, vector and the payload
|
||||
ids.append(i)
|
||||
@@ -96,17 +96,17 @@ Add all created embeddings, along with their ids and payloads into the `COCO` co
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
from qdrant_client.http.models import Batch, VectorParams, Distance
|
||||
from qdrant_client.models import Batch, VectorParams, Distance
|
||||
|
||||
qdrant_client = qdrant_client.QdrantClient()
|
||||
qdrant_client.recreate_collection(
|
||||
client = qdrant_client.QdrantClient()
|
||||
client.recreate_collection(
|
||||
collection_name="COCO",
|
||||
vectors_config=VectorParams(
|
||||
size=len(vectors[0]),
|
||||
distance=Distance.COSINE,
|
||||
),
|
||||
)
|
||||
qdrant_client.upsert(
|
||||
client.upsert(
|
||||
collection_name="COCO",
|
||||
points=Batch(
|
||||
ids=ids,
|
||||
@@ -126,7 +126,7 @@ text queries and reverse image search. Assume you want to find images similar to
|
||||
With the following code snippet create its vector embedding and then perform the lookup in Qdrant:
|
||||
|
||||
```python
|
||||
async with AsyncCliet(token=aa_token) as client:
|
||||
async with AsyncCliet(token=aa_token) as aa_client:
|
||||
prompt = ImagePrompt.from_file("query.jpg")
|
||||
prompt = Prompt.from_image(prompt)
|
||||
|
||||
@@ -136,9 +136,9 @@ async with AsyncCliet(token=aa_token) as client:
|
||||
"compress_to_size": 128,
|
||||
}
|
||||
query_request = SemanticEmbeddingRequest(**query_params)
|
||||
query_response = await client.semantic_embed(request=query_request, model=model)
|
||||
query_response = await aa_client.semantic_embed(request=query_request, model=model)
|
||||
|
||||
results = qdrant.search(
|
||||
results = client.search(
|
||||
collection_name="COCO",
|
||||
query_vector=query_response.embedding,
|
||||
limit=3,
|
||||
@@ -156,16 +156,16 @@ and Spanish. Your search is not only multimodal, but also multilingual, without
|
||||
```python
|
||||
text = "Surfing"
|
||||
|
||||
async with AsyncClient(token=aa_token) as client:
|
||||
async with AsyncClient(token=aa_token) as aa_client:
|
||||
query_params = {
|
||||
"prompt": Prompt.from_text(text),
|
||||
"representation": SemanticRepresentation.Symmetric,
|
||||
"compres_to_size": 128,
|
||||
}
|
||||
query_request = SemanticEmbeddingRequest(**query_params)
|
||||
query_response = await client.semantic_embed(request=query_request, model=model)
|
||||
query_response = await aa_client.semantic_embed(request=query_request, model=model)
|
||||
|
||||
results = qdrant.search(
|
||||
results = client.search(
|
||||
collection_name="COCO",
|
||||
query_vector=query_response.embedding,
|
||||
limit=3,
|
||||
|
||||
@@ -7,7 +7,7 @@ weight: 14
|
||||
|
||||
Asynchronous programming is being broadly adopted in the Python ecosystem. Tools such as FastAPI [have embraced this new
|
||||
paradigm](https://fastapi.tiangolo.com/async/), but it is also becoming a standard for ML models served as SaaS. For example, the Cohere SDK
|
||||
[provides an async client](https://cohere-sdk.readthedocs.io/en/latest/cohere.html#asyncclient) next to its synchronous counterpart.
|
||||
[provides an async client](https://github.com/cohere-ai/cohere-python/blob/856a4c3bd29e7a75fa66154b8ac9fcdf1e0745e0/src/cohere/client.py#L189) next to its synchronous counterpart.
|
||||
|
||||
Databases are often launched as separate services and are accessed via a network. All the interactions with them are IO-bound and can
|
||||
be performed asynchronously so as not to waste time actively waiting for a server response. In Python, this is achieved by
|
||||
|
||||
@@ -37,7 +37,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -78,7 +78,7 @@ PATCH /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.update_collection(
|
||||
collection_name="{collection_name}",
|
||||
@@ -138,7 +138,7 @@ PUT /collections/{collection_name}
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
|
||||
@@ -0,0 +1,304 @@
|
||||
---
|
||||
title: Implement Cohere RAG connector
|
||||
weight: 24
|
||||
---
|
||||
|
||||
# Implement custom connector for Cohere RAG
|
||||
|
||||
| Time: 45 min | Level: Intermediate | | |
|
||||
|--------------|---------------------|-|----|
|
||||
|
||||
The usual approach to implementing Retrieval Augmented Generation requires users to build their prompts with the
|
||||
relevant context the LLM may rely on, and manually sending them to the model. Cohere is quite unique here, as their
|
||||
models can now speak to the external tools and extract meaningful data on their own. You can virtually connect any data
|
||||
source and let the Cohere LLM know how to access it. Obviously, vector search goes well with LLMs, and enabling semantic
|
||||
search over your data is a typical case.
|
||||
|
||||
Cohere RAG has lots of interesting features, such as inline citations, which help you to refer to the specific parts of
|
||||
the documents used to generate the response.
|
||||
|
||||

|
||||
|
||||
*Source: https://docs.cohere.com/docs/retrieval-augmented-generation-rag*
|
||||
|
||||
The connectors have to implement a specific interface and expose the data source as HTTP REST API. Cohere documentation
|
||||
[describes a general process of creating a connector](https://docs.cohere.com/docs/creating-and-deploying-a-connector).
|
||||
This tutorial guides you step by step on building such a service around Qdrant.
|
||||
|
||||
## Qdrant connector
|
||||
|
||||
You probably already have some collections you would like to bring to the LLM. Maybe your pipeline was set up using some
|
||||
of the popular libraries such as Langchain, Llama Index, or Haystack. Cohere connectors may implement even more complex
|
||||
logic, e.g. hybrid search. In our case, we are going to start with a fresh Qdrant collection, index data using Cohere
|
||||
Embed v3, build the connector, and finally connect it with the [Command-R model](https://txt.cohere.com/command-r/).
|
||||
|
||||
### Building the collection
|
||||
|
||||
First things first, let's build a collection and configure it for the Cohere `embed-multilingual-v3.0` model. It
|
||||
produces 1024-dimensional embeddings, and we can choose any of the distance metrics available in Qdrant. Our connector
|
||||
will act as a personal assistant of a software engineer, and it will expose our notes to suggest the priorities or
|
||||
actions to perform.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
"https://my-cluster.cloud.qdrant.io:6333",
|
||||
api_key="my-api-key",
|
||||
)
|
||||
client.create_collection(
|
||||
collection_name="personal-notes",
|
||||
vectors_config=models.VectorParams(
|
||||
size=1024,
|
||||
distance=models.Distance.DOT,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Our notes will be represented as simple JSON objects with a `title` and `text` of the specific note. The embeddings will
|
||||
be created from the `text` field only.
|
||||
|
||||
```python
|
||||
notes = [
|
||||
{
|
||||
"title": "Project Alpha Review",
|
||||
"text": "Review the current progress of Project Alpha, focusing on the integration of the new API. Check for any compatibility issues with the existing system and document the steps needed to resolve them. Schedule a meeting with the development team to discuss the timeline and any potential roadblocks."
|
||||
},
|
||||
{
|
||||
"title": "Learning Path Update",
|
||||
"text": "Update the learning path document with the latest courses on React and Node.js from Pluralsight. Schedule at least 2 hours weekly to dedicate to these courses. Aim to complete the React course by the end of the month and the Node.js course by mid-next month."
|
||||
},
|
||||
{
|
||||
"title": "Weekly Team Meeting Agenda",
|
||||
"text": "Prepare the agenda for the weekly team meeting. Include the following topics: project updates, review of the sprint backlog, discussion on the new feature requests, and a brainstorming session for improving remote work practices. Send out the agenda and the Zoom link by Thursday afternoon."
|
||||
},
|
||||
{
|
||||
"title": "Code Review Process Improvement",
|
||||
"text": "Analyze the current code review process to identify inefficiencies. Consider adopting a new tool that integrates with our version control system. Explore options such as GitHub Actions for automating parts of the process. Draft a proposal with recommendations and share it with the team for feedback."
|
||||
},
|
||||
{
|
||||
"title": "Cloud Migration Strategy",
|
||||
"text": "Draft a plan for migrating our current on-premise infrastructure to the cloud. The plan should cover the selection of a cloud provider, cost analysis, and a phased migration approach. Identify critical applications for the first phase and any potential risks or challenges. Schedule a meeting with the IT department to discuss the plan."
|
||||
},
|
||||
{
|
||||
"title": "Quarterly Goals Review",
|
||||
"text": "Review the progress towards the quarterly goals. Update the documentation to reflect any completed objectives and outline steps for any remaining goals. Schedule individual meetings with team members to discuss their contributions and any support they might need to achieve their targets."
|
||||
},
|
||||
{
|
||||
"title": "Personal Development Plan",
|
||||
"text": "Reflect on the past quarter's achievements and areas for improvement. Update the personal development plan to include new technical skills to learn, certifications to pursue, and networking events to attend. Set realistic timelines and check-in points to monitor progress."
|
||||
},
|
||||
{
|
||||
"title": "End-of-Year Performance Reviews",
|
||||
"text": "Start preparing for the end-of-year performance reviews. Collect feedback from peers and managers, review project contributions, and document achievements. Consider areas for improvement and set goals for the next year. Schedule preliminary discussions with each team member to gather their self-assessments."
|
||||
},
|
||||
{
|
||||
"title": "Technology Stack Evaluation",
|
||||
"text": "Conduct an evaluation of our current technology stack to identify any outdated technologies or tools that could be replaced for better performance and productivity. Research emerging technologies that might benefit our projects. Prepare a report with findings and recommendations to present to the management team."
|
||||
},
|
||||
{
|
||||
"title": "Team Building Event Planning",
|
||||
"text": "Plan a team-building event for the next quarter. Consider activities that can be done remotely, such as virtual escape rooms or online game nights. Survey the team for their preferences and availability. Draft a budget proposal for the event and submit it for approval."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
Storing the embeddings along with the metadata is fairly simple.
|
||||
|
||||
```python
|
||||
import cohere
|
||||
import uuid
|
||||
|
||||
cohere_client = cohere.Client(api_key="my-cohere-api-key")
|
||||
|
||||
response = cohere_client.embed(
|
||||
texts=[
|
||||
note.get("text")
|
||||
for note in notes
|
||||
],
|
||||
model="embed-multilingual-v3.0",
|
||||
input_type="search_document",
|
||||
)
|
||||
|
||||
client.upload_points(
|
||||
collection_name="personal-notes",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=uuid.uuid4().hex,
|
||||
vector=embedding,
|
||||
payload=note,
|
||||
)
|
||||
for note, embedding in zip(notes, response.embeddings)
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Our collection is now ready to be searched over. In the real world, the set of notes would be changing over time, so the
|
||||
ingestion process won't be as straightforward. This data is not yet exposed to the LLM, but we will build the connector
|
||||
in the next step.
|
||||
|
||||
### Connector web service
|
||||
|
||||
[FastAPI](https://fastapi.tiangolo.com/) is a modern web framework and perfect a choice for a simple HTTP API. We are
|
||||
going to use it for the purposes of our connector. There will be just one endpoint, as required by the model. It will
|
||||
accept POST requests at the `/search` path. There is a single `query` parameter required. Let's define a corresponding
|
||||
model.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
|
||||
class SearchQuery(BaseModel):
|
||||
query: str
|
||||
```
|
||||
|
||||
RAG connector does not have to return the documents in any specific format. There are [some good practices to follow](https://docs.cohere.com/docs/creating-and-deploying-a-connector#configure-the-connection-between-the-connector-and-the-chat-api),
|
||||
but Cohere models are quite flexible here. Results just have to be returned as JSON, with a list of objects in a
|
||||
`results` property of the output. We will use the same document structure as we did for the Qdrant payloads, so there
|
||||
is no conversion required. That requires two additional models to be created.
|
||||
|
||||
```python
|
||||
from typing import List
|
||||
|
||||
class Document(BaseModel):
|
||||
title: str
|
||||
text: str
|
||||
|
||||
class SearchResults(BaseModel):
|
||||
results: List[Document]
|
||||
```
|
||||
|
||||
Once our model classes are ready, we can implement the logic that will get the query and provide the notes that are
|
||||
relevant to it. Please note the LLM is not going to define the number of documents to be returned. That's completely
|
||||
up to you how many of them you want to bring to the context.
|
||||
|
||||
There are two services we need to interact with - Qdrant server and Cohere API. FastAPI has a concept of a [dependency
|
||||
injection](https://fastapi.tiangolo.com/tutorial/dependencies/#dependencies), and we will use it to provide both
|
||||
clients into the implementation.
|
||||
|
||||
In case of queries, we need to set the `input_type` to `search_query` in the calls to Cohere API.
|
||||
|
||||
```python
|
||||
from fastapi import FastAPI, Depends
|
||||
from typing import Annotated
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
def client() -> QdrantClient:
|
||||
return QdrantClient(config.QDRANT_URL, api_key=config.QDRANT_API_KEY)
|
||||
|
||||
def cohere_client() -> cohere.Client:
|
||||
return cohere.Client(api_key=config.COHERE_API_KEY)
|
||||
|
||||
@app.post("/search")
|
||||
def search(
|
||||
query: SearchQuery,
|
||||
client: Annotated[QdrantClient, Depends(client)],
|
||||
cohere_client: Annotated[cohere.Client, Depends(cohere_client)],
|
||||
) -> SearchResults:
|
||||
response = cohere_client.embed(
|
||||
texts=[query.query],
|
||||
model="embed-multilingual-v3.0",
|
||||
input_type="search_query",
|
||||
)
|
||||
results = client.search(
|
||||
collection_name="personal-notes",
|
||||
query_vector=response.embeddings[0],
|
||||
limit=2,
|
||||
)
|
||||
return SearchResults(
|
||||
results=[
|
||||
Document(**point.payload)
|
||||
for point in results
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Our app might be launched locally for the development purposes, given we have the `uvicorn` server installed:
|
||||
|
||||
```shell
|
||||
uvicorn main:app
|
||||
```
|
||||
|
||||
FastAPI exposes an interactive documentation at `http://localhost:8000/docs`, where we can test our endpoint. The
|
||||
`/search` endpoint is available there.
|
||||
|
||||

|
||||
|
||||
We can interact with it and check the documents that will be returned for a specific query. For example, we want to know
|
||||
recall what we are supposed to do regarding the infrastructure for your projects.
|
||||
|
||||
```shell
|
||||
curl -X "POST" \
|
||||
-H "Content-type: application/json" \
|
||||
-d '{"query": "Is there anything I have to do regarding the project infrastructure?"}' \
|
||||
"http://localhost:8000/search"
|
||||
```
|
||||
|
||||
The output should look like following:
|
||||
|
||||
```json
|
||||
{
|
||||
"results": [
|
||||
{
|
||||
"title": "Cloud Migration Strategy",
|
||||
"text": "Draft a plan for migrating our current on-premise infrastructure to the cloud. The plan should cover the selection of a cloud provider, cost analysis, and a phased migration approach. Identify critical applications for the first phase and any potential risks or challenges. Schedule a meeting with the IT department to discuss the plan."
|
||||
},
|
||||
{
|
||||
"title": "Project Alpha Review",
|
||||
"text": "Review the current progress of Project Alpha, focusing on the integration of the new API. Check for any compatibility issues with the existing system and document the steps needed to resolve them. Schedule a meeting with the development team to discuss the timeline and any potential roadblocks."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Connecting to Command-R
|
||||
|
||||
Our web service is implemented, yet running only on our local machine. It has to be exposed to the public before
|
||||
Command-R can interact with it. For a quick experiment, it might be enough to set up tunneling using services such as
|
||||
[ngrok](https://ngrok.com/). We won't cover all the details in the tutorial, but their
|
||||
[Quickstart](https://ngrok.com/docs/guides/getting-started/) is a great resource describing the process step-by-step.
|
||||
Alternatively, you can also deploy the service with a public URL.
|
||||
|
||||
Once it's done, we can create the connector first, and then tell the model to use it, while interacting through the chat
|
||||
API. Creating a connector is a single call to Cohere client:
|
||||
|
||||
```python
|
||||
connector_response = cohere_client.connectors.create(
|
||||
name="personal-notes",
|
||||
url="https:/this-is-my-domain.app/search",
|
||||
)
|
||||
```
|
||||
|
||||
The `connector_response.connector` will be a descriptor, with `id` being one of the attributes. We'll use this
|
||||
identifier for our interactions like this:
|
||||
|
||||
```python
|
||||
response = cohere_client.chat(
|
||||
message=(
|
||||
"Is there anything I have to do regarding the project infrastructure? "
|
||||
"Please mention the tasks briefly."
|
||||
),
|
||||
connectors=[
|
||||
cohere.ChatConnector(id=connector_response.connector.id)
|
||||
],
|
||||
model="command-r",
|
||||
)
|
||||
```
|
||||
|
||||
We changed the `model` to `command-r`, as this is currently the best Cohere model available to public. The
|
||||
`response.text` is the output of the model:
|
||||
|
||||
```text
|
||||
Here are some of the tasks related to project infrastructure that you might have to perform:
|
||||
- You need to draft a plan for migrating your on-premise infrastructure to the cloud and come up with a plan for the selection of a cloud provider, cost analysis, and a gradual migration approach.
|
||||
- It's important to evaluate your current technology stack to identify any outdated technologies. You should also research emerging technologies and the benefits they could bring to your projects.
|
||||
```
|
||||
|
||||
You only need to create a specific connector once! Please do not call `cohere_client.connectors.create` for every single
|
||||
message you send to the `chat` method.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
We have built a Cohere RAG connector that integrates with your existing knowledge base stored in Qdrant. We covered just
|
||||
the basic flow, but in real world scenarios, you should also consider e.g. [building the authentication
|
||||
system](https://docs.cohere.com/docs/connector-authentication) to prevent unauthorized access.
|
||||
@@ -75,9 +75,9 @@ We used the streaming mode, so the dataset is not loaded into memory. Instead, w
|
||||
|
||||
```python
|
||||
for payload in dataset:
|
||||
id = payload.pop("id")
|
||||
id_ = payload.pop("id")
|
||||
vector = payload.pop("vector")
|
||||
print(id, vector, payload)
|
||||
print(id_, vector, payload)
|
||||
```
|
||||
|
||||
A single payload looks like this:
|
||||
@@ -114,10 +114,10 @@ Calculating the embeddings is usually a bottleneck of the vector search pipeline
|
||||
```python
|
||||
ids, vectors, payloads = [], [], []
|
||||
for payload in dataset:
|
||||
id = payload.pop("id")
|
||||
id_ = payload.pop("id")
|
||||
vector = payload.pop("vector")
|
||||
|
||||
ids.append(id)
|
||||
ids.append(id_)
|
||||
vectors.append(vector)
|
||||
payloads.append(payload)
|
||||
|
||||
|
||||
@@ -106,7 +106,7 @@ Now you need to write a script to upload all startup data and vectors into the s
|
||||
# Import client library
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
qdrant_client = QdrantClient("http://localhost:6333")
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
```
|
||||
|
||||
3. Select model to encode your data.
|
||||
@@ -114,16 +114,16 @@ qdrant_client = QdrantClient("http://localhost:6333")
|
||||
You will be using a pre-trained model called `sentence-transformers/all-MiniLM-L6-v2`.
|
||||
|
||||
```python
|
||||
qdrant_client.set_model("sentence-transformers/all-MiniLM-L6-v2")
|
||||
client.set_model("sentence-transformers/all-MiniLM-L6-v2")
|
||||
```
|
||||
|
||||
|
||||
4. Related vectors need to be added to a collection. Create a new collection for your startup vectors.
|
||||
|
||||
```python
|
||||
qdrant_client.recreate_collection(
|
||||
client.recreate_collection(
|
||||
collection_name="startups",
|
||||
vectors_config=qdrant_client.get_fastembed_vector_params(),
|
||||
vectors_config=client.get_fastembed_vector_params(),
|
||||
)
|
||||
```
|
||||
|
||||
|
||||
@@ -146,13 +146,13 @@ Now you need to write a script to upload all startup data and vectors into the s
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.models import VectorParams, Distance
|
||||
|
||||
qdrant_client = QdrantClient("http://localhost:6333")
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
```
|
||||
|
||||
3. Related vectors need to be added to a collection. Create a new collection for your startup vectors.
|
||||
|
||||
```python
|
||||
qdrant_client.recreate_collection(
|
||||
client.recreate_collection(
|
||||
collection_name="startups",
|
||||
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
|
||||
)
|
||||
@@ -186,7 +186,7 @@ vectors = np.load("./startup_vectors.npy")
|
||||
5. Upload the data
|
||||
|
||||
```python
|
||||
qdrant_client.upload_collection(
|
||||
client.upload_collection(
|
||||
collection_name="startups",
|
||||
vectors=vectors,
|
||||
payload=payload,
|
||||
|
||||
@@ -105,10 +105,10 @@ after receiving the response from the `upsert` endpoint. **As long as the indexi
|
||||
the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed.
|
||||
|
||||
```python
|
||||
client.upload_records(
|
||||
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
records=[
|
||||
models.Record(
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=item["id"],
|
||||
vector=item["vector"],
|
||||
payload=item,
|
||||
|
||||
@@ -147,7 +147,7 @@ documents = [
|
||||
You need to tell Qdrant where to store embeddings. This is a basic demo, so your local computer will use its memory as temporary storage.
|
||||
|
||||
```python
|
||||
qdrant = QdrantClient(":memory:")
|
||||
client = QdrantClient(":memory:")
|
||||
```
|
||||
|
||||
## 4. Create a collection
|
||||
@@ -155,7 +155,7 @@ qdrant = QdrantClient(":memory:")
|
||||
All data in Qdrant is organized by collections. In this case, you are storing books, so we are calling it `my_books`.
|
||||
|
||||
```python
|
||||
qdrant.recreate_collection(
|
||||
client.recreate_collection(
|
||||
collection_name="my_books",
|
||||
vectors_config=models.VectorParams(
|
||||
size=encoder.get_sentence_embedding_dimension(), # Vector size is defined by used model
|
||||
@@ -176,7 +176,7 @@ qdrant.recreate_collection(
|
||||
Tell the database to upload `documents` to the `my_books` collection. This will give each record an id and a payload. The payload is just the metadata from the dataset.
|
||||
|
||||
```python
|
||||
qdrant.upload_points(
|
||||
client.upload_points(
|
||||
collection_name="my_books",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
@@ -192,7 +192,7 @@ qdrant.upload_points(
|
||||
Now that the data is stored in Qdrant, you can ask it questions and receive semantically relevant results.
|
||||
|
||||
```python
|
||||
hits = qdrant.search(
|
||||
hits = client.search(
|
||||
collection_name="my_books",
|
||||
query_vector=encoder.encode("alien invasion").tolist(),
|
||||
limit=3,
|
||||
@@ -216,7 +216,7 @@ The search engine shows three of the most likely responses that have to do with
|
||||
How about the most recent book from the early 2000s?
|
||||
|
||||
```python
|
||||
hits = qdrant.search(
|
||||
hits = client.search(
|
||||
collection_name="my_books",
|
||||
query_vector=encoder.encode("alien invasion").tolist(),
|
||||
query_filter=models.Filter(
|
||||
|
||||
@@ -7,4 +7,3 @@ sitemapExclude: True
|
||||
Current advances in NLP can reduce the retinue work of customer service by up to 80 percent.
|
||||
No more answering the same questions over and over again. A chatbot will do that, and people can focus on complex problems.
|
||||
But not only automated answering, it is also possible to control the quality of the department and automatically identify flaws in conversations.
|
||||
Read more about the "[Sentence Embeddings for Customer Support](https://blog.floydhub.com/automate-customer-support-part-one/)" case study.
|
||||
|
After Width: | Height: | Size: 89 KiB |
|
After Width: | Height: | Size: 58 KiB |
|
After Width: | Height: | Size: 89 KiB |
|
After Width: | Height: | Size: 58 KiB |
|
After Width: | Height: | Size: 7.0 KiB |
|
After Width: | Height: | Size: 4.2 KiB |
|
After Width: | Height: | Size: 44 KiB |
|
After Width: | Height: | Size: 89 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 14 KiB |
|
After Width: | Height: | Size: 25 KiB |
|
After Width: | Height: | Size: 15 KiB |
|
After Width: | Height: | Size: 106 KiB |
|
After Width: | Height: | Size: 73 KiB |
|
After Width: | Height: | Size: 45 KiB |
|
After Width: | Height: | Size: 618 KiB |
|
After Width: | Height: | Size: 988 KiB |
|
After Width: | Height: | Size: 28 KiB |
|
After Width: | Height: | Size: 21 KiB |
|
After Width: | Height: | Size: 178 KiB |
|
After Width: | Height: | Size: 116 KiB |
|
After Width: | Height: | Size: 84 KiB |
|
After Width: | Height: | Size: 543 KiB |
|
After Width: | Height: | Size: 415 KiB |
|
After Width: | Height: | Size: 26 KiB |
|
After Width: | Height: | Size: 21 KiB |
|
After Width: | Height: | Size: 158 KiB |
|
After Width: | Height: | Size: 104 KiB |
|
After Width: | Height: | Size: 74 KiB |
|
After Width: | Height: | Size: 21 KiB |
|
After Width: | Height: | Size: 16 KiB |
|
After Width: | Height: | Size: 137 KiB |
|
After Width: | Height: | Size: 88 KiB |
|
After Width: | Height: | Size: 61 KiB |
|
After Width: | Height: | Size: 23 KiB |