mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-01 00:48:32 +02:00
Merge branch 'master' of https://github.com/qdrant/landing_page into colpali-13
This commit is contained in:
@@ -19,7 +19,7 @@ features:
|
||||
description: By combining dense vector embeddings with sparse vectors e.g. BM25, Qdrant powers semantic search to deliver context-aware results, transcending traditional keyword search by understanding the deeper meaning of data.
|
||||
link:
|
||||
text: Learn More
|
||||
url: /documentation/tutorials/hybrid-search-fastembed/
|
||||
url: /documentation/beginner-tutorials/hybrid-search-fastembed/
|
||||
- id: 2
|
||||
icon:
|
||||
src: /icons/outline/selection-blue.svg
|
||||
@@ -28,12 +28,12 @@ features:
|
||||
description: Qdrant's capability extends to multi-modal search, indexing and retrieving various data forms (text, images, audio) once vectorized, facilitating a comprehensive search experience.
|
||||
link:
|
||||
text: View Tutorial
|
||||
url: /documentation/tutorials/aleph-alpha-search/
|
||||
url: /documentation/tutorials/multimodal-search-fastembed/
|
||||
- id: 3
|
||||
icon:
|
||||
src: /icons/outline/filter-blue.svg
|
||||
alt: Filter
|
||||
title: Single Stage filtering that Works
|
||||
title: Single Stage filtering That Works
|
||||
description: Qdrant enhances search speeds and control and context understanding through filtering on any nested entry in our payload. Unique architecture allows Qdrant to avoid expensive pre-filtering and post-filtering stages, making search faster and accurate.
|
||||
link:
|
||||
text: Learn More
|
||||
|
||||
@@ -14,11 +14,11 @@ features:
|
||||
image:
|
||||
src: /img/advanced-search-use-cases/multimodal-semantic-search.svg
|
||||
alt: Multimodal Semantic Search
|
||||
title: Multimodal Semantic Search with Aleph Alpha
|
||||
title: Multimodal Semantic Search with FastEmbed
|
||||
description: This tutorial shows you how to run a proper multimodal semantic search system with a few lines of code, without the need to annotate the data or train your networks.
|
||||
link:
|
||||
text: View Tutorial
|
||||
url: /documentation/examples/aleph-alpha-search/
|
||||
url: /documentation/tutorials/multimodal-search-fastembed/
|
||||
- id: 2
|
||||
image:
|
||||
src: /img/advanced-search-use-cases/simple-neural-search.svg
|
||||
@@ -27,7 +27,7 @@ features:
|
||||
description: This tutorial shows you how to build and deploy your own neural search service.
|
||||
link:
|
||||
text: View Tutorial
|
||||
url: /documentation/tutorials/neural-search/
|
||||
url: /documentation/beginner-tutorials/neural-search/
|
||||
- id: 3
|
||||
image:
|
||||
src: /img/advanced-search-use-cases/image-classification.svg
|
||||
@@ -45,16 +45,16 @@ features:
|
||||
description: Build a semantic search engine for science fiction books in 5 mins.
|
||||
link:
|
||||
text: View Tutorial
|
||||
url: /documentation/tutorials/search-beginners/
|
||||
url: /documentation/beginner-tutorials/search-beginners/
|
||||
- id: 5
|
||||
image:
|
||||
src: /img/advanced-search-use-cases/hybrid-search-service-fastembed.svg
|
||||
alt: Create a Hybrid Search Service with Fastembed
|
||||
title: Create a Hybrid Search Service with Fastembed
|
||||
description: This tutorial guides you through building and deploying your own hybrid search service using Fastembed.
|
||||
alt: Create a Hybrid Search Service with FastEmbed
|
||||
title: Create a Hybrid Search Service with FastEmbed
|
||||
description: This tutorial guides you through building and deploying your own hybrid search service using FastEmbed.
|
||||
link:
|
||||
text: View Tutorial
|
||||
url: /documentation/tutorials/hybrid-search-fastembed/
|
||||
url: /documentation/beginner-tutorials/hybrid-search-fastembed/
|
||||
sitemapExclude: true
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
title: "AI Agents"
|
||||
description: "AI Agents"
|
||||
build:
|
||||
render: always
|
||||
cascade:
|
||||
- build:
|
||||
list: local
|
||||
publishResources: false
|
||||
render: never
|
||||
---
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
image:
|
||||
src: /img/ai-agents-dashboard-cloud.png
|
||||
alt: Dashboard cloud
|
||||
title: AI Agents with Qdrant
|
||||
description: AI agents powered by Qdrant leverage advanced vector search to access and retrieve high-dimensional data in real-time, enabling intelligent, Agentic-RAG driven, multi-step decision-making across dynamic environments.
|
||||
cases:
|
||||
- id: 0
|
||||
title: Multimodal Data Handling
|
||||
description: Qdrant enables AI agents to process and retrieve high-dimensional vectors from diverse data types (text, images, audio), supporting more comprehensive decision-making in multimodal environments.
|
||||
- id: 1
|
||||
title: Adaptive Learning
|
||||
description: Qdrant supports continuous learning by enabling efficient vector retrieval and updates, allowing agents to learn and evolve based on real-time interactions and new data points.
|
||||
featuresTitle: Qdrant equips AI agents to adapt, learn, and collaborate efficiently.
|
||||
features:
|
||||
- id: 0
|
||||
icon:
|
||||
src: /icons/outline/precision-blue.svg
|
||||
alt: Precision
|
||||
title: Contextual Precision
|
||||
description: Qdrant’s hybrid search combines semantic vector search, lexical search, and metadata filtering, enabling AI Agents to retrieve highly relevant and contextually precise information. This enhances decision-making by allowing agents to leverage both meaning-based and keyword-based strategies, ensuring accuracy and relevance for complex queries in dynamic environments.
|
||||
link:
|
||||
text: Hybrid Search
|
||||
url: /articles/hybrid-search/
|
||||
- id: 1
|
||||
icon:
|
||||
src: /icons/outline/multitenancy-blue.svg
|
||||
alt: Multitenancy
|
||||
title: Multi-Agent Systems
|
||||
description: Qdrant’s scalability and multitenancy ensures that multiple agents can collaborate in distributed systems, enabling seamless coordination and communication - key for Agentic RAG workflows.
|
||||
link:
|
||||
text: Multitenancy
|
||||
url: /articles/multitenancy/
|
||||
- id: 2
|
||||
icon:
|
||||
src: /icons/outline/time-blue.svg
|
||||
alt: Time
|
||||
title: Real Time Decision Making
|
||||
description: Qdrant’s real-time, advanced vector search enables AI agents to act instantly on live data, which is crucial for time-sensitive, autonomous decision-making.
|
||||
link:
|
||||
text: HNSW
|
||||
url: /articles/filtrable-hnsw/
|
||||
- id: 3
|
||||
icon:
|
||||
src: /icons/outline/server-rack-blue.svg
|
||||
alt: Server rack
|
||||
title: Optimized CPU Performance for Embedding Processing
|
||||
description: Qdrant’s architecture is optimized for high-throughput embedding processing, minimizing CPU load and preventing performance bottlenecks. This enables AI agents in Agentic RAG workflows to execute complex, multi-step tasks efficiently, ensuring smooth operation even at scale.
|
||||
link:
|
||||
text: Distributed Deployment
|
||||
url: /documentation/guides/distributed_deployment/
|
||||
- id: 4
|
||||
icon:
|
||||
src: /icons/outline/speedometer-blue.svg
|
||||
alt: Speedometer
|
||||
title: Semantic Cache for Rapid Query Handling
|
||||
description: Qdrant enhances AI agent efficiency with semantic caching, which preserves results of queries based on semantic equivalence rather than exact matches. This method reduces query processing times and system load by reusing previously computed answers, essential for high-throughput AI applications.
|
||||
link:
|
||||
text: Semantic Cache
|
||||
url: /articles/semantic-cache-ai-data-retrieval/
|
||||
sitemapExclude: true
|
||||
---
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
title: Building AI agents?
|
||||
description: Apply for the Qdrant for Startups program to access a 20% discount to Qdrant Cloud, our managed cloud service, perks from Hugging Face, LlamaIndex, and Airbyte, and much more.
|
||||
button:
|
||||
url: /qdrant-for-startups/
|
||||
text: Apply Now
|
||||
image:
|
||||
src: /img/ai-agent.svg
|
||||
alt: AI agent
|
||||
sitemapExclude: true
|
||||
---
|
||||
@@ -0,0 +1,15 @@
|
||||
---
|
||||
title: AI Agents
|
||||
description: Unlock the full potential of your AI agents with Qdrant’s powerful vector search and scalable infrastructure, allowing them to handle complex tasks, adapt in real time, and drive smarter, data-driven outcomes across any environment.
|
||||
startFree:
|
||||
text: Get Started
|
||||
url: https://cloud.qdrant.io/
|
||||
learnMore:
|
||||
text: Learn More
|
||||
url: "#ai-agents"
|
||||
image:
|
||||
src: /img/vectors/vector-4.svg
|
||||
alt: AI agents chat
|
||||
sitemapExclude: true
|
||||
---
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
---
|
||||
title: Qdrant integrates with the leading AI Agent frameworks
|
||||
integrations:
|
||||
- id: 0
|
||||
icon:
|
||||
src: /img/integrations/integration-lang-graph.svg
|
||||
alt: LangGraph logo
|
||||
title: LangGraph
|
||||
description: Framework for managing multi-LLM workflows with structured graphs for AI agents.
|
||||
url: /documentation/frameworks/langgraph/
|
||||
- id: 1
|
||||
icon:
|
||||
src: /img/integrations/integration-open-ai.svg
|
||||
alt: Open AI logo
|
||||
title: Swarm
|
||||
description: Decentralized platform enabling collaboration among AI agents for task completion.
|
||||
url: /documentation/frameworks/swarm/
|
||||
- id: 2
|
||||
icon:
|
||||
src: /img/integrations/integration-crew-ai.svg
|
||||
alt: Crew AI logo
|
||||
title: CrewAI
|
||||
description: Team-based AI agent collaboration system orchestrating multi-agent workflows efficiently.
|
||||
url: /documentation/frameworks/crewai/
|
||||
- id: 3
|
||||
icon:
|
||||
src: /img/integrations/integration-auto-gen.svg
|
||||
alt: AutoGen logo
|
||||
title: AutoGen
|
||||
description: Automation tool for generating AI agent workflows and automating complex tasks.
|
||||
url: /documentation/frameworks/autogen/
|
||||
sitemapExclude: true
|
||||
---
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
title: Learn how to get started with Qdrant for your AI Agent use case
|
||||
features:
|
||||
- id: 0
|
||||
image:
|
||||
src: /img/ai-agents-use-cases/comparing-ai-agent-frameworks.svg
|
||||
srcMobile: /img/ai-agents-use-cases/comparing-ai-agent-frameworks.svg
|
||||
alt: Comparing AI agent frameworks
|
||||
title: "Comparing AI Agent Frameworks: LangGraph, CrewAI, Swarm, and AutoGen"
|
||||
description: This guide offers a comparison of key AI agent frameworks, highlighting their strengths and ideal use cases for developers.
|
||||
link:
|
||||
text: View Guide
|
||||
url: /articles/agentic-rag/
|
||||
- id: 1
|
||||
image:
|
||||
src: /img/ai-agents-use-cases/open-ai-agents.svg
|
||||
srcMobile: /img/ai-agents-use-cases/open-ai-agents.svg
|
||||
alt: Open AI agents
|
||||
title: Building OpenAI Swarm Agents with Qdrant
|
||||
description: Learn how to build OpenAI Swarm agents using Qdrant for fast, scalable vector search and real-time actions.
|
||||
link:
|
||||
text: Read the Docs
|
||||
url: /documentation/frameworks/swarm/
|
||||
- id: 2
|
||||
image:
|
||||
src: /img/ai-agents-use-cases/ai-scheduler.svg
|
||||
srcMobile: /img/ai-agents-use-cases/ai-scheduler.svg
|
||||
alt: AI scheduler
|
||||
title: Building an AI Scheduler with Zoom, CrewAI, and Qdrant
|
||||
description: Learn how to build an AI meeting scheduler with Zoom, LlamaIndex, and Qdrant, featuring a hands-on RAG recommendation engine code sample.
|
||||
link:
|
||||
text: View Tutorial
|
||||
url: /documentation/agentic-rag-crewai-zoom/
|
||||
caseStudy:
|
||||
logo:
|
||||
src: /img/ai-agents-use-cases/customer-logo.svg
|
||||
alt: Logo
|
||||
title: "QA.tech Case Study: AI Agents for Web Testing"
|
||||
description: QA.tech enhanced web app testing by deploying AI agents that mimic user interactions. To handle high-speed actions and make real-time decisions, they integrated Qdrant for scalable vector search, allowing for faster and more efficient data proc...
|
||||
link:
|
||||
text: Read Case Study
|
||||
url: /blog/case-study-qatech/
|
||||
image:
|
||||
src: /img/ai-agents-use-cases/case-study.png
|
||||
alt: Preview
|
||||
sitemapExclude: true
|
||||
---
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
---
|
||||
tag:
|
||||
title: On-demand Webinar
|
||||
icon:
|
||||
src: /icons/outline/training-purple.svg
|
||||
alt: Training
|
||||
title: Building AI Agents for personalized recommendations with Qdrant and n8n
|
||||
description: Learn in this video how to build an AI-powered recommendation system using Qdrant and n8n. It demonstrates how an AI agent retrieves data from Qdrant's vector database and leverages a large language model (LLM) to generate personalized recommendations based on user inputs.
|
||||
link:
|
||||
text: Watch Now
|
||||
url: https://www.youtube.com/watch?v=O5mT8M7rqQQ
|
||||
youtube: |
|
||||
<iframe src="https://www.youtube.com/embed/O5mT8M7rqQQ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
sitemapExclude: true
|
||||
---
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
tag:
|
||||
title: On-demand Webinar
|
||||
icon:
|
||||
src: /icons/outline/training-purple.svg
|
||||
alt: Training
|
||||
title: Building Agents with LlamaIndex & Qdrant
|
||||
description: Ready to build more advanced AI agents? Watch this webinar to learn how to use LlamaIndex and Qdrant to create intelligent agents capable of handling complex, multi-modal queries in RAG-enabled systems.
|
||||
link:
|
||||
text: Watch Now
|
||||
url: https://www.youtube.com/watch?v=3mWIFsooibQ
|
||||
image:
|
||||
src: /img/ai-agents-webinar.svg
|
||||
alt: AI agents webinar
|
||||
sitemapExclude: true
|
||||
---
|
||||
|
||||
@@ -0,0 +1,455 @@
|
||||
---
|
||||
title: "What is Agentic RAG? Building Agents with Qdrant"
|
||||
short_description: "Will agentic RAG replace linear RAG? Learn how to build agents with Qdrant and which framework is best for your use case."
|
||||
description: "Agents are a new paradigm in AI, and they are changing how we build RAG systems. Learn how to build agents with Qdrant and which framework to choose."
|
||||
preview_dir: /articles_data/agentic-rag/preview
|
||||
social_preview_image: /articles_data/agentic-rag/social-preview.png
|
||||
weight: -150
|
||||
author: Kacper Łukawski
|
||||
author_link: https://www.kacperlukawski.com
|
||||
date: 2024-11-22T00:00:00.000Z
|
||||
---
|
||||
|
||||
Standard [Retrieval Augmented Generation](/articles/what-is-rag-in-ai/) follows a predictable, linear path: receive
|
||||
a query, retrieve relevant documents, and generate a response. In many cases that might be enough to solve a particular
|
||||
problem. In the worst case scenario, your LLM will just decide to not answer the question, because the context does not
|
||||
provide enough information.
|
||||
|
||||

|
||||
|
||||
On the other hand, we have agents. These systems are given more freedom to act, and can take multiple non-linear steps
|
||||
to achieve a certain goal. There isn't a single definition of what an agent is, but in general, it is an application
|
||||
that uses LLM and usually some tools to communicate with the outside world. LLMs are used as decision-makers which
|
||||
decide what action to take next. Actions can be anything, but they are usually well-defined and limited to a certain
|
||||
set of possibilities. One of these actions might be to query a vector database, like Qdrant, to retrieve relevant
|
||||
documents, if the context is not enough to make a decision. However, RAG is just a single tool in the agent's arsenal.
|
||||
|
||||

|
||||
|
||||
## Agentic RAG: Combining RAG with Agents
|
||||
|
||||
Since the agent definition is vague, the concept of **Agentic RAG** is also not well-defined. In general, it refers to
|
||||
the combination of RAG with agents. This allows the agent to use external knowledge sources to make decisions, and
|
||||
primarily to decide when the external knowledge is needed. We can describe a system as Agentic RAG if it breaks the
|
||||
linear flow of a standard RAG system, and gives the agent the ability to take multiple steps to achieve a goal.
|
||||
|
||||
A simple router that chooses a path to follow is often described as the simplest form of an agent. Such a system has
|
||||
multiple paths with conditions describing when to take a certain path. In the context of Agentic RAG, the agent can
|
||||
decide to query a vector database if the context is not enough to answer, or skip the query if it's enough, or when the
|
||||
question refers to common knowledge. Alternatively, there might be multiple collections storing different kinds of
|
||||
information, and the agent can decide which collection to query based on the context. The key factor is that the
|
||||
decision of choosing a path is made by the LLM, which is the core of the agent. A routing agent never comes back to the
|
||||
previous step, so it's ultimately just a conditional decision-making system.
|
||||
|
||||

|
||||
|
||||
However, routing is just the beginning. Agents can be much more complex, and extreme forms of agents can have complete
|
||||
freedom to act. In such cases, the agent is given a set of tools and can autonomously decide which ones to use, how to
|
||||
use them, and in which order. LLMs are asked to plan and execute actions, and the agent can take multiple steps to
|
||||
achieve a goal, including taking steps back if needed. Such a system does not have to follow a DAG structure (Directed
|
||||
Acyclic Graph), and can have loops that help to self-correct the decisions made in the past. An agentic RAG system
|
||||
built in that manner can have tools not only to query a vector database, but also to play with the query, summarize the
|
||||
results, or even generate new data to answer the question. Options are endless, but there are some common patterns
|
||||
that can be observed in the wild.
|
||||
|
||||

|
||||
|
||||
### Solving Information Retrieval Problems with LLMs
|
||||
|
||||
Generally speaking, tools exposed in an agentic RAG system are used to solve information retrieval problems which are
|
||||
not new to the search community. LLMs have changed how we approach these problems, but the core of the problem remains
|
||||
the same. What kind of tools you can consider using in an agentic RAG? Here are some examples:
|
||||
|
||||
- **Querying a vector database** - the most common tool used in agentic RAG systems. It allows the agent to retrieve
|
||||
relevant documents based on the query.
|
||||
- **Query expansion** - a tool that can be used to improve the query. It can be used to add synonyms, correct typos, or
|
||||
even to generate new queries based on the original one.
|
||||

|
||||
- **Extracting filters** - vector search alone is sometimes not enough. In many cases, you might want to narrow down
|
||||
the results based on specific parameters. This extraction process can automatically identify relevant conditions from
|
||||
the query. Otherwise, your users would have to manually define these search constraints.
|
||||

|
||||
- **Quality judgement** - knowing the quality of the results for given query can be used to decide whether they are good
|
||||
enough to answer, or if the agent should take another step to improve them somehow. Alternatively it can also admit
|
||||
the failure to provide good response.
|
||||

|
||||
|
||||
These are just some of the examples, but the list is not exhaustive. For example, your LLM could possibly play with
|
||||
Qdrant search parameters or choose different methods to query it. An example? If your users are searching using some
|
||||
specific keywords, you may prefer sparse vectors to dense vectors, as they are more efficient in such cases. In that
|
||||
case you have to arm your agent with tools to decide when to use sparse vectors and when to use dense vectors. Agent
|
||||
aware of the collection structure can make such decisions easily.
|
||||
|
||||
Each of these tools might be a separate agent on its own, and multi-agent systems are not uncommon. In such cases,
|
||||
agents can communicate with each other, and one agent can decide to use another agent to solve a particular problem.
|
||||
Pretty useful component of an agentic RAG is also a human in the loop, which can be used to correct the agent's
|
||||
decisions, or steer it in the right direction.
|
||||
|
||||
## Where are Agents Used?
|
||||
|
||||
Agents are an interesting concept, but since they heavily rely on LLMs, they are not applicable to all problems. Using
|
||||
Large Language Models is expensive and tend to be slow, what in many cases, it's not worth the cost. Standard RAG
|
||||
involves just a single call to the LLM, and the response is generated in a predictable way. Agents, on the other hand,
|
||||
can take multiple steps, and the latency experienced by the user adds up. In many cases, it's not acceptable.
|
||||
Agentic RAG is probably not that widely applicable in ecommerce search, where the user expects a quick response, but
|
||||
might be fine for customer support, where the user is willing to wait a bit longer for a better answer.
|
||||
|
||||
## Which Framework is Best?
|
||||
|
||||
There are lots of frameworks available to build agents, and choosing the best one is not easy. It depends on your
|
||||
existing stack or the tools you are familiar with. Some of the most popular LLM libraries have already drifted towards
|
||||
the agent paradigm, and they are offering tools to build them. There are, however, some tools built primarily for
|
||||
agents development, so let's focus on them.
|
||||
|
||||
### LangGraph
|
||||
|
||||
Developed by the LangChain team, LangGraph seems like a natural extension for those who already use LangChain for
|
||||
building their RAG systems, and would like to start with agentic RAG.
|
||||
|
||||
Surprisingly, LangGraph has nothing to do with Large Language Models on its own. It's a framework for building
|
||||
graph-based applications in which each **node** is a step of the workflow. Each node takes an application **state** as
|
||||
an input, and produces a modified state as an output. The state is then passed to the next node, and so on. **Edges**
|
||||
between the nodes might be conditional what makes branching possible. Contrary to some DAG-based tool (i.e. Apache
|
||||
Airflow), LangGraph allows for loops in the graph, which makes it possible to implement cyclic workflows, so an agent
|
||||
can achieve self-reflection and self-correction. Theoretically, LangGraph can be used to build any kind of applications
|
||||
in a graph-based manner, not only LLM agents.
|
||||
|
||||
Some of the strengths of LangGraph include:
|
||||
|
||||
- **Persistence** - the state of the workflow graph is stored as a checkpoint. That happens at each so-called super-step
|
||||
(which is a single sequential node of a graph). It enables replying certain steps of the workflow, fault-tolerance,
|
||||
and including human-in-the-loop interactions. This mechanism also acts as a **short-term memory**, accessible in a
|
||||
context of a particular workflow execution.
|
||||
- **Long-term memory** - LangGraph also has a concept of memories that are shared between different workflow runs.
|
||||
However, this mechanism has to explicitly handled by our nodes. **Qdrant with its semantic search capabilities is
|
||||
often used as a long-term memory layer**.
|
||||
- **Multi-agent support** - while there is no separate concept of multi-agent systems in LangGraph, it's possible to
|
||||
create such an architecture by building a graph that includes multiple agents and some kind of supervisor that
|
||||
makes a decision which agent to use in a given situation. If a node might be anything, then it might be another agent
|
||||
as well.
|
||||
|
||||
Some other interesting features of LangGraph include the ability to visualize the graph, automate the retries of failed
|
||||
steps, and include human-in-the-loop interactions.
|
||||
|
||||
A minimal example of an agentic RAG could improve the user query, e.g. by fixing typos, expanding it with synonyms, or
|
||||
even generating a new query based on the original one. The agent could then retrieve documents from a vector database
|
||||
based on the improved query, and generate a response. The LangGraph app implementing this approach could look like this:
|
||||
|
||||
```python
|
||||
from typing import Sequence
|
||||
from typing_extensions import TypedDict, Annotated
|
||||
from langchain_core.messages import BaseMessage
|
||||
from langgraph.constants import START, END
|
||||
from langgraph.graph import add_messages, StateGraph
|
||||
|
||||
|
||||
class AgentState(TypedDict):
|
||||
# The state of the agent includes at least the messages exchanged between the agent(s)
|
||||
# and the user. It is, however, possible to include other information in the state, as
|
||||
# it depends on the specific agent.
|
||||
messages: Annotated[Sequence[BaseMessage], add_messages]
|
||||
|
||||
|
||||
def improve_query(state: AgentState):
|
||||
...
|
||||
|
||||
def retrieve_documents(state: AgentState):
|
||||
...
|
||||
|
||||
def generate_response(state: AgentState):
|
||||
...
|
||||
|
||||
# Building a graph requires defining nodes and building the flow between them with edges.
|
||||
builder = StateGraph(AgentState)
|
||||
|
||||
builder.add_node("improve_query", improve_query)
|
||||
builder.add_node("retrieve_documents", retrieve_documents)
|
||||
builder.add_node("generate_response", generate_response)
|
||||
|
||||
builder.add_edge(START, "improve_query")
|
||||
builder.add_edge("improve_query", "retrieve_documents")
|
||||
builder.add_edge("retrieve_documents", "generate_response")
|
||||
builder.add_edge("generate_response", END)
|
||||
|
||||
# Compiling the graph performs some checks and prepares the graph for execution.
|
||||
compiled_graph = builder.compile()
|
||||
|
||||
# Compiled graph might be invoked with the initial state to start.
|
||||
compiled_graph.invoke({
|
||||
"messages": [
|
||||
("user", "Why Qdrant is the best vector database out there?"),
|
||||
]
|
||||
})
|
||||
```
|
||||
|
||||
Each node of the process is just a Python function that does certain operation. You can call an LLM of your choice
|
||||
inside of them, if you want to, but there is no assumption about the messages being created by any AI. **LangGraph
|
||||
rather acts as a runtime that launches these functions in a specific order, and passes the state between them**. While
|
||||
[LangGraph](https://www.langchain.com/langgraph) integrates well with the LangChain ecosystem, it can be used
|
||||
independently. For teams looking for additional support and features, there's also a commercial offering called
|
||||
LangGraph Platform. The framework is available for both Python and JavaScript environments, making it possible to be
|
||||
used in different tech stacks.
|
||||
|
||||
### CrewAI
|
||||
|
||||
CrewAI is another popular choice for building agents, including agentic RAG. It's a high-level framework that assumes
|
||||
there are some LLM-based agents working together to achieve a common goal. That's where the "crew" in CrewAI comes from.
|
||||
CrewAI is designed with multi-agent systems in mind. Contrary to LangGraph, the developer does not create a graph of
|
||||
processing, but defines agents and their roles within the crew.
|
||||
|
||||
Some of the key concepts of CrewAI include:
|
||||
|
||||
- **Agent** - a unit that has a specific role and goal, controlled by an LLM. It can optionally use some external tools
|
||||
to communicate with the outside world, but generally steered by prompt we provide to the LLM.
|
||||
- **Process** - currently either sequential or hierarchical. It defines how the task will be executed by the agents.
|
||||
In a sequential process, agents are executed one after another, while in a hierarchical process, agent is selected
|
||||
by the manager agent, which is responsible for making decisions about which agent to use in a given situation.
|
||||
- **Roles and goals** - each agent has a certain role within the crew, and the goal it should aim to achieve. These are
|
||||
set when we define an agent and are used to make decisions about which agent to use in a given situation.
|
||||
- **Memory** - an extensive memory system consists of short-term memory, long-term memory, entity memory, and contextual
|
||||
memory that combines the other three. There is also user memory for preferences and personalization. **This is where
|
||||
Qdrant comes into play, as it might be used as a long-term memory layer.**
|
||||
|
||||
CrewAI provides a rich set of tools integrated into the framework. That may be a huge advantage for those who want to
|
||||
combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however brining your own tools is
|
||||
not a big deal, as CrewAI is designed to be extensible.
|
||||
|
||||
A simple agentic RAG application implemented in CrewAI could look like this:
|
||||
|
||||
```python
|
||||
from crewai import Crew, Agent, Task
|
||||
from crewai.memory.entity.entity_memory import EntityMemory
|
||||
from crewai.memory.short_term.short_term_memory import ShortTermMemory
|
||||
from crewai.memory.storage.rag_storage import RAGStorage
|
||||
|
||||
class QdrantStorage(RAGStorage):
|
||||
...
|
||||
|
||||
response_generator_agent = Agent(
|
||||
role="Generate response based on the conversation",
|
||||
goal="Provide the best response, or admit when the response is not available.",
|
||||
backstory=(
|
||||
"I am a response generator agent. I generate "
|
||||
"responses based on the conversation."
|
||||
),
|
||||
verbose=True,
|
||||
)
|
||||
|
||||
query_reformulation_agent = Agent(
|
||||
role="Reformulate the query",
|
||||
goal="Rewrite the query to get better results. Fix typos, grammar, word choice, etc.",
|
||||
backstory=(
|
||||
"I am a query reformulation agent. I reformulate the "
|
||||
"query to get better results."
|
||||
),
|
||||
verbose=True,
|
||||
)
|
||||
|
||||
task = Task(
|
||||
description="Let me know why Qdrant is the best vector database out there.",
|
||||
expected_output="3 bullet points",
|
||||
agent=response_generator_agent,
|
||||
)
|
||||
|
||||
crew = Crew(
|
||||
agents=[response_generator_agent, query_reformulation_agent],
|
||||
tasks=[task],
|
||||
memory=True,
|
||||
entity_memory=EntityMemory(storage=QdrantStorage("entity")),
|
||||
short_term_memory=ShortTermMemory(storage=QdrantStorage("short-term")),
|
||||
)
|
||||
crew.kickoff()
|
||||
```
|
||||
|
||||
*Disclaimer: QdrantStorage is not a part of the CrewAI framework, but it's taken from the Qdrant documentation on [how
|
||||
to integrate Qdrant with CrewAI](https://qdrant.tech/documentation/frameworks/crewai/).*
|
||||
|
||||
Although it's not a technical advantage, CrewAI has a [great documentation](https://docs.crewai.com/introduction). The
|
||||
framework is available for Python, and it's easy to get started with it. CrewAI also has a commercial offering, CrewAI
|
||||
Enterprise, which provides a platform for building and deploying agents at scale.
|
||||
|
||||
### AutoGen
|
||||
|
||||
AutoGen emphasizes multi-agent architectures as a fundamental design principle. The framework requires at least two
|
||||
agents in any system to really call an application agentic - typically an assistant and a user proxy exchange messages
|
||||
to achieve a common goal. Sequential chat with more than two agents is also supported, as well as group chat and nested
|
||||
chat for internal dialogue. However, AutoGen does not assume there is a structured state that is passed between the
|
||||
agents, and the chat conversation is the only way to communicate between them.
|
||||
|
||||
There are many interesting concepts in the framework, some of them even quite unique:
|
||||
|
||||
- **Tools/functions** - external components that can be used by agents to communicate with the outside world. They are
|
||||
defined as Python callables, and can be used for any external interaction we want to allow the agent to do. Type
|
||||
annotations are used to define the input and output of the tools, and Pydantic models are supported for more complex
|
||||
type schema. AutoGen supports only OpenAI-compatible tool call API for the time being.
|
||||
- **Code executors** - built-in code executors include local command, Docker command, and Jupyter. An agent can write
|
||||
and launch code, so theoretically the agents can do anything that can be done in Python. None of the other frameworks
|
||||
made code generation and execution that prominent. Code execution being the first-class citizen in AutoGen is an
|
||||
interesting concept.
|
||||
|
||||
Each AutoGen agent uses at least one of the components: human-in-the-loop, code executor, tool executor, or LLM.
|
||||
A simple agentic RAG, based on the conversation of two agents which can retrieve documents from a vector database,
|
||||
or improve the query, could look like this:
|
||||
|
||||
```python
|
||||
from os import environ
|
||||
|
||||
from autogen import ConversableAgent
|
||||
from autogen.agentchat.contrib.retrieve_user_proxy_agent import RetrieveUserProxyAgent
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(...)
|
||||
|
||||
response_generator_agent = ConversableAgent(
|
||||
name="response_generator_agent",
|
||||
system_message=(
|
||||
"You answer user questions based solely on the provided context. You ask to retrieve relevant documents for "
|
||||
"your query, or reformulate the query, if it is incorrect in some way."
|
||||
),
|
||||
description="A response generator agent that can answer your queries.",
|
||||
llm_config={"config_list": [{"model": "gpt-4", "api_key": environ.get("OPENAI_API_KEY")}]},
|
||||
human_input_mode="NEVER",
|
||||
)
|
||||
|
||||
user_proxy = RetrieveUserProxyAgent(
|
||||
name="retrieval_user",
|
||||
llm_config={"config_list": [{"model": "gpt-4", "api_key": environ.get("OPENAI_API_KEY")}]},
|
||||
human_input_mode="NEVER",
|
||||
retrieve_config={
|
||||
"task": "qa",
|
||||
"chunk_token_size": 2000,
|
||||
"vector_db": "qdrant",
|
||||
"db_config": {"client": client},
|
||||
"get_or_create": True,
|
||||
"overwrite": True,
|
||||
},
|
||||
)
|
||||
|
||||
result = user_proxy.initiate_chat(
|
||||
response_generator_agent,
|
||||
message=user_proxy.message_generator,
|
||||
problem="Why Qdrant is the best vector database out there?",
|
||||
max_turns=10,
|
||||
)
|
||||
```
|
||||
|
||||
For those new to agent development, AutoGen offers AutoGen Studio, a low-code interface for prototyping agents. While
|
||||
not intended for production use, it significantly lowers the barrier to entry for experimenting with agent
|
||||
architectures.
|
||||
|
||||

|
||||
|
||||
It's worth noting that AutoGen is currently undergoing significant updates, with version 0.4.x in development
|
||||
introducing substantial API changes compared to the stable 0.2.x release. While the framework currently has limited
|
||||
built-in persistence and state management capabilities, these features may evolve in future releases.
|
||||
|
||||
### OpenAI Swarm
|
||||
|
||||
Unliked the other frameworks described in this article, OpenAI Swarm is an educational project, and it's not ready for
|
||||
production use. It's worth mentioning, though, as it's pretty lightweight and easy to get started with. OpenAI Swarm
|
||||
is an experimental framework for orchestrating multi-agent workflows that focuses on agent coordination through direct
|
||||
handoffs rather than complex orchestration patterns.
|
||||
|
||||
With that setup, **agents** are just exchanging messages in a chat, optionally calling some Python functions to
|
||||
communicate with external services, or handing off the conversation to another agent, if the other one seems to be more
|
||||
suitable to answer the question. Each agent has a certain role, defined by the instructions we have to define.
|
||||
We have to decide which LLM will a particular agent use, and a set of functions it can call. For example, **a retrieval
|
||||
agent could use a vector database to retrieve documents**, and return the results to the next agent. That means, there
|
||||
should be a function that performs the semantic search on its behalf, but the model will decide how the query should
|
||||
look like.
|
||||
|
||||
Here is how a similar agentic RAG application, implemented in OpenAI Swarm, could look like:
|
||||
|
||||
```python
|
||||
from swarm import Swarm, Agent
|
||||
|
||||
client = Swarm()
|
||||
|
||||
def retrieve_documents(query: str) -> list[str]:
|
||||
"""
|
||||
Retrieve documents based on the query.
|
||||
"""
|
||||
...
|
||||
|
||||
def transfer_to_query_improve_agent():
|
||||
return query_improve_agent
|
||||
|
||||
query_improve_agent = Agent(
|
||||
name="Query Improve Agent",
|
||||
instructions=(
|
||||
"You are a search expert that takes user queries and improves them to get better results. You fix typos and "
|
||||
"extend queries with synonyms, if needed. You never ask the user for more information."
|
||||
),
|
||||
)
|
||||
|
||||
response_generation_agent = Agent(
|
||||
name="Response Generation Agent",
|
||||
instructions=(
|
||||
"You take the whole conversation and generate a final response based on the chat history. "
|
||||
"If you don't have enough information, you can retrieve the documents from the knowledge base or "
|
||||
"reformulate the query by transferring to other agent. You never ask the user for more information. "
|
||||
"You have to always be the last participant of each conversation."
|
||||
),
|
||||
functions=[retrieve_documents, transfer_to_query_improve_agent],
|
||||
)
|
||||
|
||||
response = client.run(
|
||||
agent=response_generation_agent,
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Why Qdrant is the best vector database out there?"
|
||||
}
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
Even though we don't explicitly define the graph of processing, the agents can still decide to hand off the processing
|
||||
to a different agent. There is no concept of a state, so everything relies on the messages exchanged between different
|
||||
components.
|
||||
|
||||
OpenAI Swarm does not focus on integration with external tools, and **if you would like to integrate semantic search
|
||||
with Qdrant, you would have to implement it fully yourself**. Obviously, the library is tightly coupled with OpenAI
|
||||
models, and while using some other ones is possible, it requires some additional work like setting up proxy that will
|
||||
adjust the interface to OpenAI API.
|
||||
|
||||
### The winner?
|
||||
|
||||
Choosing the best framework for your agentic RAG system depends on your existing stack, team expertise, and the
|
||||
specific requirements of your project. All the described tools are strong contenders, and they are developed at rapid
|
||||
pace. It's worth keeping an eye on all of them, as they are likely to evolve and improve over time. Eventually, you
|
||||
should be able to build the same processes with any of them, but some of them may be more suitable in a specific
|
||||
ecosystem of the tools you want your agent to interact with.
|
||||
|
||||
There are, however, some important factors to consider when choosing a framework for your agentic RAG system:
|
||||
|
||||
- **Human-in-the-loop** - even though we aim to build autonomous agents, it's often important to include the feedback
|
||||
from the human, so our agents cannot perform malicious actions.
|
||||
- **Observability** - how easy it is to debug the system, and how easy it is to understand what's happening inside.
|
||||
Especially important, since we are dealing with lots of LLM prompts.
|
||||
|
||||
Still, choosing the right toolkit depends on the state of your project, and the specific requirements you have. If you
|
||||
want to integrate your agent with number of external tools, CrewAI might be the best choice, as the set of
|
||||
out-of-the-box integrations is the biggest. However, LangGraph integrates well with LangChain, so if you are familiar
|
||||
with that ecosystem, it may suit you better.
|
||||
|
||||
All the frameworks have different approaches to building agents, so it's worth experimenting with all of them to see
|
||||
which one fits your needs the best. LangGraph and CrewAI are more mature and have more features, while AutoGen and
|
||||
OpenAI Swarm are more lightweight and more experimental. However, **none of the existing frameworks solves all the
|
||||
mentioned Information Retrieval problems**, so you still have to build your own tools to fill the gaps.
|
||||
|
||||
## Building Agentic RAG with Qdrant
|
||||
|
||||
No matter which framework you choose, Qdrant is a great tool to build agentic RAG systems. Please check out [our
|
||||
integrations](/documentation/frameworks/) to choose the best one for your use case and preferences. The easiest way to
|
||||
start using Qdrant is to use our managed service, [Qdrant Cloud](https://cloud.qdrant.io). A free 1GB cluster is
|
||||
available for free, so you can start building your agentic RAG system in minutes.
|
||||
|
||||
### Further Reading
|
||||
|
||||
See how Qdrant integrates with:
|
||||
|
||||
- [Autogen](https://qdrant.tech/documentation/frameworks/autogen/)
|
||||
- [CrewAI](https://qdrant.tech/documentation/frameworks/crewai/)
|
||||
- [LangGraph](https://qdrant.tech/documentation/frameworks/langgraph/)
|
||||
- [Swarm](https://qdrant.tech/documentation/frameworks/swarm/)
|
||||
@@ -305,7 +305,7 @@ query_text = "best programming language for beginners?"
|
||||
model_bm42 = SparseTextEmbedding(model_name="Qdrant/bm42-all-minilm-l6-v2-attentions")
|
||||
model_jina = TextEmbedding(model_name="jinaai/jina-embeddings-v2-base-en")
|
||||
|
||||
sparse_embedding = list(embedding_model.query_embed(query_text))[0]
|
||||
sparse_embedding = list(model_bm42.query_embed(query_text))[0]
|
||||
dense_embedding = list(model_jina.query_embed(query_text))[0]
|
||||
|
||||
client.query_points(
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
---
|
||||
title: Qdrant Summer of Code 2024 - ONNX Cross Encoders in Python
|
||||
short_description: QSoC 2024 ONNX Cross Encoders in Python
|
||||
description: A summary of my work and experience at Qdrant Summer of Code 2024.
|
||||
preview_dir: /articles_data/cross-encoder-integration-gsoc/preview
|
||||
small_preview_image: /articles_data/cross-encoder-integration-gsoc/icon.svg
|
||||
social_preview_image: /articles_data/cross-encoder-integration-gsoc/preview/social_preview.jpg
|
||||
weight: -212
|
||||
author: Huong (Celine) Hoang
|
||||
author_link: https://www.linkedin.com/in/celine-h-hoang/
|
||||
date: 2024-10-14T08:00:00+03:00
|
||||
draft: false
|
||||
keywords:
|
||||
- cross-encoder
|
||||
- reranking
|
||||
- fastembed
|
||||
- qsoc'24
|
||||
---
|
||||
|
||||
## Introduction
|
||||
|
||||
Hi everyone! I’m Huong (Celine) Hoang, and I’m thrilled to share my experience working at Qdrant this summer as part of their Summer of Code 2024 program. During my internship, I worked on integrating cross-encoders into the FastEmbed library for re-ranking tasks. This enhancement widened the capabilities of the Qdrant ecosystem, enabling developers to build more context-aware search applications, such as question-answering systems, using Qdrant's suite of libraries.
|
||||
|
||||
This project was both technically challenging and rewarding, pushing me to grow my skills in handling large-scale ONNX (Open Neural Network Exchange) model integrations, tokenization, and more. Let me take you through the journey, the lessons learned, and where things are headed next.
|
||||
|
||||
## Project Overview
|
||||
|
||||
Qdrant is well known for its vector search capabilities, but my task was to go one step further — introducing cross-encoders for re-ranking. Traditionally, the FastEmbed library would generate embeddings, but cross-encoders don’t do that. Instead, they provide a list of scores based on how well a query matches a list of documents. This kind of re-ranking is critical when you want to refine search results and bring the most relevant answers to the top.
|
||||
|
||||
The project revolved around creating a new input-output scheme: text data to scores. For this, I designed a family of classes to support ONNX models. Some of the key models I worked with included Xenova/ms-marco-MiniLM-L-6-v2, Xenova/ms-marco-MiniLM-L-12-v2, and BAAI/bge-reranker, all designed for re-ranking tasks.
|
||||
|
||||
An important point to mention is that FastEmbed is a minimalistic library: it doesn’t have heavy dependencies like PyTorch or TensorFlow, and as a result, it is lightweight, occupying far less storage space.
|
||||
|
||||
Below is a diagram that represents the overall workflow for this project, detailing the key steps from user interaction to the final output validation:
|
||||
|
||||
{{< figure src="/articles_data/cross-encoder-integration-gsoc/rerank-workflow.png" caption="Search workflow with reranking" alt="Search workflow with reranking" >}}
|
||||
|
||||
## Technical Challenges
|
||||
|
||||
### 1. Building a New Input-Output Scheme
|
||||
|
||||
FastEmbed already had support for embeddings, but re-ranking with cross-encoders meant building a completely new family of classes. These models accept a query and a set of documents, then return a list of relevance scores. For that, I created the base classes like `TextCrossEncoderBase` and `OnnxCrossEncoder`, taking inspiration from existing text embedding models.
|
||||
|
||||
One thing I had to ensure was that the new class hierarchy was user-friendly. Users should be able to work with cross-encoders without needing to know the complexities of the underlying models. For instance, they should be able to just write:
|
||||
|
||||
```python
|
||||
from fastembed.rerank.cross_encoder import TextCrossEncoder
|
||||
|
||||
encoder = TextCrossEncoder(model_name="Xenova/ms-marco-MiniLM-L-6-v2")
|
||||
scores = encoder.rerank(query, documents)
|
||||
```
|
||||
|
||||
Meanwhile, behind the scenes, we manage all the model loading, tokenization, and scoring.
|
||||
|
||||
### 2. Handling Tokenization for Cross-Encoders
|
||||
|
||||
Cross-encoders require careful tokenization because they need to distinguish between the query and the documents. This is done using token type IDs, which help the model differentiate between the two. To implement this, I configured the tokenizer to handle pairs of inputs—concatenating the query with each document and assigning token types accordingly.
|
||||
|
||||
Efficient tokenization is critical to ensure the performance of the models, and I optimized it specifically for ONNX models.
|
||||
|
||||
### 3. Model Loading and Integration
|
||||
|
||||
One of the most rewarding parts of the project was integrating the ONNX models into the FastEmbed library. ONNX models need to be loaded into a runtime environment that efficiently manages the computations.
|
||||
|
||||
While PyTorch is a common framework for these types of tasks, FastEmbed exclusively supports ONNX models, making it both lightweight and efficient. I focused on extensive testing to ensure that the ONNX models performed equivalently to their PyTorch counterparts, ensuring users could trust the results.
|
||||
|
||||
I added support for batching as well, allowing users to re-rank large sets of documents without compromising speed.
|
||||
|
||||
### 4. Debugging and Code Reviews
|
||||
|
||||
During the project, I encountered a number of challenges, including issues with model configurations, tokenizers, and test cases. With the help of my mentor, George Panchuk, I was able to resolve these issues and improve my understanding of best practices, particularly around code readability, maintainability, and style.
|
||||
|
||||
One notable lesson was the importance of keeping the code organized and maintainable, with a strong focus on readability. This included properly structuring modules and ensuring the entire codebase followed a clear, consistent style.
|
||||
|
||||
### 5. Testing and Validation
|
||||
To ensure the accuracy and performance of the models, I conducted extensive testing. I compared the output of ONNX models with their PyTorch counterparts, ensuring the conversion to ONNX was correct. A key part of this process was rigorous testing to verify the outputs and identify potential issues, such as incorrect conversions or bugs in our implementation.
|
||||
|
||||
For instance, a test to validate the model's output was structured as follows:
|
||||
```python
|
||||
def test_rerank():
|
||||
is_ci = os.getenv("CI")
|
||||
|
||||
for model_desc in TextCrossEncoder.list_supported_models():
|
||||
if not is_ci and model_desc["size_in_GB"] > 1:
|
||||
continue
|
||||
|
||||
model_name = model_desc["model"]
|
||||
model = TextCrossEncoder(model_name=model_name)
|
||||
|
||||
query = "What is the capital of France?"
|
||||
documents = ["Paris is the capital of France.", "Berlin is the capital of Germany."]
|
||||
scores = np.array(model.rerank(query, documents))
|
||||
|
||||
canonical_scores = CANONICAL_SCORE_VALUES[model_name]
|
||||
assert np.allclose(
|
||||
scores, canonical_scores, atol=1e-3
|
||||
), f"Model: {model_name}, Scores: {scores}, Expected: {canonical_scores}"
|
||||
```
|
||||
|
||||
The `CANONICAL_SCORE_VALUES` were retrieved directly from the result of applying the original PyTorch models to the same input
|
||||
|
||||
## Outcomes and Future Improvements
|
||||
|
||||
By the end of my project, I successfully added cross-encoders to the FastEmbed library, allowing users to re-rank search results based on relevance scores. This enhancement opens up new possibilities for applications that rely on contextual ranking, such as search engines and recommendation systems.
|
||||
This functionality will be available as of FastEmbed `0.4.0`.
|
||||
|
||||
Some areas for future improvements include:
|
||||
- Expanding Model Support: We could add more cross-encoder models, especially from the sentence transformers library, to give users more options.
|
||||
- Parallelization: Optimizing batch processing to handle even larger datasets could further improve performance.
|
||||
- Custom Tokenization: For models with non-standard tokenization, like BAAI/bge-reranker, more specific tokenizer configurations could be added.
|
||||
|
||||
## Overall Experience and Wrapping Up
|
||||
|
||||
Looking back, this internship has been an incredibly valuable experience. I’ve grown not only as a developer but also as someone who can take on complex projects and see them through from start to finish. The Qdrant team has been so supportive, especially during the debugging and review stages. I’ve learned so much about model integration, ONNX, and how to build tools that are user-friendly and scalable.
|
||||
|
||||
One key takeaway for me is the importance of understanding the user experience. It’s not just about getting the models to work but making sure they are easy to use and integrate into real-world applications. This experience has solidified my passion for building solutions that truly make an impact, and I’m excited to continue working on projects like this in the future.
|
||||
|
||||
Thank you for taking the time to read about my journey with Qdrant and the FastEmbed library. I’m excited to see how this work will continue to improve search experiences for users!
|
||||
@@ -0,0 +1,623 @@
|
||||
---
|
||||
title: "Modern Sparse Neural Retrieval: From Theory to Practice"
|
||||
short_description: ""
|
||||
description: "A comprehensive guide to modern sparse neural retrievers: COIL, TILDEv2, SPLADE, and more. Find out how they work and learn how to use them effectively."
|
||||
preview_dir: /articles_data/modern-sparse-neural-retrieval/preview
|
||||
social_preview_image: /articles_data/modern-sparse-neural-retrieval/social-preview.png
|
||||
weight: -213
|
||||
author: Evgeniya Sukhodolskaya
|
||||
date: 2024-10-23T00:00:00.000Z
|
||||
tags:
|
||||
- sparse retriever
|
||||
- sparse retrieval
|
||||
- splade
|
||||
- bm25
|
||||
---
|
||||
|
||||
Finding enough time to study all the modern solutions while keeping your production running is rarely feasible.
|
||||
Dense retrievers, hybrid retrievers, late interaction… How do they work, and where do they fit best?
|
||||
If only we could compare retrievers as easily as products on Amazon!
|
||||
|
||||
We explored the most popular modern sparse neural retrieval models and broke them down for you.
|
||||
By the end of this article, you’ll have a clear understanding of the current landscape in sparse neural retrieval and how to navigate through complex, math-heavy research papers with sky-high NDCG scores without getting overwhelmed.
|
||||
|
||||
[The first part](#sparse-neural-retrieval-evolution) of this article is theoretical, comparing different approaches used in
|
||||
modern sparse neural retrieval.\
|
||||
[The second part](#splade-in-qdrant) is more practical, showing how the best model in modern sparse neural retrieval, `SPLADE++`,
|
||||
can be used in Qdrant and recommendations on when to choose sparse neural retrieval for your solutions.
|
||||
|
||||
## Sparse Neural Retrieval: As If Keyword-Based Retrievers Understood Meaning
|
||||
|
||||
**Keyword-based (lexical) retrievers** like BM25 provide a good explainability.
|
||||
If a document matches a query, it’s easy to understand why: query terms are present in the document,
|
||||
and if these are rare terms, they are more important for retrieval.
|
||||
|
||||

|
||||
|
||||
With their mechanism of exact term matching, they are super fast at retrieval.
|
||||
A simple **inverted index**, which maps back from a term to a list of documents where this term occurs, saves time on checking millions of documents.
|
||||
|
||||

|
||||
|
||||
Lexical retrievers are still a strong baseline in retrieval tasks.
|
||||
However, by design, they’re unable to bridge **vocabulary** and **semantic mismatch** gaps.
|
||||
Imagine searching for a “*tasty cheese*” in an online store and not having a chance to get “*Gouda*” or “*Brie*” in your shopping basket.
|
||||
|
||||
**Dense retrievers**, based on machine learning models which encode documents and queries in dense vector representations,
|
||||
are capable of breaching this gap and finding you “*a piece of Gouda*”.
|
||||
|
||||

|
||||
|
||||
However, explainability here suffers: why is this query representation close to this document representation?
|
||||
Why, searching for “*cheese*”, we’re also offered “*mouse traps*”? What does each number in this vector representation mean?
|
||||
Which one of them is capturing the cheesiness?
|
||||
|
||||
Without a solid understanding, balancing result quality and resource consumption becomes challenging.
|
||||
Since, hypothetically, any document could match a query, relying on an inverted index with exact matching isn’t feasible.
|
||||
This doesn’t mean dense retrievers are inherently slower. However, lexical retrieval has been around long enough to inspire several effective architectural choices, which are often worth reusing.
|
||||
|
||||
Sooner or later, there should have been somebody who would say,
|
||||
“*Wait, but what if I want something timeproof like BM25 but with semantic understanding?*”
|
||||
|
||||
## Sparse Neural Retrieval Evolution
|
||||
|
||||
Imagine searching for a “*flabbergasting murder*” story.
|
||||
”*Flabbergasting*” is a rarely used word, so a keyword-based retriever, for example, BM25, will assign huge importance to it.
|
||||
Consequently, there is a high chance that a text unrelated to any crimes but mentioning something “*flabbergasting*” will pop up in the top results.
|
||||
|
||||
What if we could instead of relying on term frequency in a document as a proxy of term’s importance as it happens in BM25,
|
||||
directly predict a term’s importance? The goal is for rare but non-impactful terms to be assigned a much smaller weight than important terms with the same frequency, while both would be equally treated in the BM25 scenario.
|
||||
|
||||
How can we determine if one term is more important than another?
|
||||
Word impact is related to its meaning, and its meaning can be derived from its context (words which surround this particular word).
|
||||
That’s how dense contextual embedding models come into the picture.
|
||||
|
||||
All the sparse retrievers are based on the idea of taking a model which produces contextual dense vector representations for terms
|
||||
and teaching it to produce sparse ones. Very often,
|
||||
[Bidirectional Encoder Representations from the Transformers (BERT)](https://huggingface.co/docs/transformers/en/model_doc/bert) is used as a
|
||||
base model, and a very simple trainable neural network is added on top of it to sparsify the representations out.
|
||||
Training this small neural network is usually done by sampling from the [MS MARCO](https://microsoft.github.io/msmarco/) dataset a query,
|
||||
relevant and irrelevant to it documents and shifting the parameters of the neural network in the direction of relevancy.
|
||||
|
||||
|
||||
### The Pioneer Of Sparse Neural Retrieval
|
||||
|
||||

|
||||
The authors of one of the first sparse retrievers, the [`Deep Contextualized Term Weighting framework (DeepCT)`](https://arxiv.org/pdf/1910.10687),
|
||||
predict an integer word’s impact value separately for each unique word in a document and a query.
|
||||
They use a linear regression model on top of the contextual representations produced by the basic BERT model, the model's output is rounded.
|
||||
|
||||
When documents are uploaded into a database, the importance of words in a document is predicted by a trained linear regression model
|
||||
and stored in the inverted index in the same way as term frequencies in BM25 retrievers.
|
||||
Then, the retrieval process is identical to the BM25 one.
|
||||
|
||||
***Why is DeepCT not a perfect solution?*** To train linear regression, the authors needed to provide the true value (**ground truth**)
|
||||
of each word’s importance so the model could “see” what the right answer should be.
|
||||
This score is hard to define in a way that it truly expresses the query-document relevancy.
|
||||
Which score should have the most relevant word to a query when this word is taken from a five-page document? The second relevant? The third?
|
||||
|
||||
### Sparse Neural Retrieval on Relevance Objective
|
||||
|
||||

|
||||
It’s much easier to define whether a document as a whole is relevant or irrelevant to a query.
|
||||
That’s why the [`DeepImpact`](https://arxiv.org/pdf/2104.12016) Sparse Neural Retriever authors directly used the relevancy between a query and a document as a training objective.
|
||||
They take BERT’s contextualized embeddings of the document’s words, transform them through a simple 2-layer neural network in a single scalar
|
||||
score and sum these scores up for each word overlapping with a query.
|
||||
The training objective is to make this score reflect the relevance between the query and the document.
|
||||
|
||||
***Why is DeepImpact not a perfect solution?***
|
||||
When converting texts into dense vector representations,
|
||||
the BERT model does not work on a word level. Sometimes, it breaks the words into parts.
|
||||
For example, the word “*vector*” will be processed by BERT as one piece, but for some words that, for example,
|
||||
BERT hasn’t seen before, it is going to cut the word in pieces
|
||||
[as “Qdrant” turns to “Q”, “#dra” and “#nt”](https://huggingface.co/spaces/Xenova/the-tokenizer-playground)
|
||||
|
||||
The DeepImpact model (like the DeepCT model) takes the first piece BERT produces for a word and discards the rest.
|
||||
However, what can one find searching for “*Q*” instead of “*Qdrant*”?
|
||||
|
||||
### Know Thine Tokenization
|
||||
|
||||

|
||||
To solve the problems of DeepImpact's architecture, the [`Term Independent Likelihood MoDEl (TILDEv2)`](https://arxiv.org/pdf/2108.08513) model generates
|
||||
sparse encodings on a level of BERT’s representations, not on words level. Aside from that, its authors use the identical architecture
|
||||
to the DeepImpact model.
|
||||
|
||||
***Why is TILDEv2 not a perfect solution?***
|
||||
A single scalar importance score value might not be enough to capture all distinct meanings of a word.
|
||||
**Homonyms** (pizza, cocktail, flower, and female name “*Margherita*”) are one of the troublemakers in information retrieval.
|
||||
|
||||
### Sparse Neural Retriever Which Understood Homonyms
|
||||
|
||||

|
||||
|
||||
If one value for the term importance score is insufficient, we could describe the term’s importance in a vector form!
|
||||
Authors of the [`COntextualized Inverted List (COIL)`](https://arxiv.org/pdf/2104.07186) model based their work on this idea.
|
||||
Instead of squeezing 768-dimensional BERT’s contextualised embeddings into one value,
|
||||
they down-project them (through the similar “relevance” training objective) to 32 dimensions.
|
||||
Moreover, not to miss a detail, they also encode the query terms as vectors.
|
||||
|
||||
For each vector representing a query token, COIL finds the closest match (using the maximum dot product) vector of the same token in a document.
|
||||
So, for example, if we are searching for “*Revolut bank \<finance institution\>*” and a document in a database has the sentence
|
||||
“*Vivid bank \<finance institution\> was moved to the bank of Amstel \<river\>*”, out of two “banks”,
|
||||
the first one will have a bigger value of a dot product with a “*bank*” in the query, and it will count towards the final score.
|
||||
The final relevancy score of a document is a sum of scores of query terms matched.
|
||||
|
||||
***Why is COIL not a perfect solution?*** This way of defining the importance score captures deeper semantics;
|
||||
more meaning comes with more values used to describe it.
|
||||
However, storing 32-dimensional vectors for every term is far more expensive,
|
||||
and an inverted index does not work as-is with this architecture.
|
||||
|
||||
### Back to the Roots
|
||||
|
||||

|
||||
[`Universal COntextualized Inverted List (UniCOIL)`](https://arxiv.org/pdf/2106.14807), made by the authors of COIL as a follow-up, goes back to producing a scalar value as the importance score
|
||||
rather than a vector, leaving unchanged all other COIL design decisions. \
|
||||
It optimizes resources consumption but the deep semantics understanding tied to COIL architecture is again lost.
|
||||
|
||||
## Did we Solve the Vocabulary Mismatch Yet?
|
||||
|
||||
With the retrieval based on the exact matching,
|
||||
however sophisticated the methods to predict term importance are, we can’t match relevant documents which have no query terms in them.
|
||||
If you’re searching for “*pizza*” in a book of recipes, you won’t find “*Margherita*”.
|
||||
|
||||
A way to solve this problem is through the so-called **document expansion**.
|
||||
Let’s append words which could be in a potential query searching for this document.
|
||||
So, the “*Margherita*” document becomes “*Margherita pizza*”. Now, exact matching on “*pizza*” will work!
|
||||
|
||||

|
||||
|
||||
There are two types of document expansion that are used in sparse neural retrieval:
|
||||
**external** (one model is responsible for expansion, another one for retrieval) and **internal** (all is done by a single model).
|
||||
|
||||
### External Document Expansion
|
||||
External document expansion uses a **generative model** (Mistral 7B, Chat-GPT, and Claude are all generative models,
|
||||
generating words based on the input text) to compose additions to documents before converting them to sparse representations
|
||||
and applying exact matching methods.
|
||||
|
||||
#### External Document Expansion with docT5query
|
||||
|
||||

|
||||
[`docT5query`](https://github.com/castorini/docTTTTTquery) is the most used document expansion model.
|
||||
It is based on the [Text-to-Text Transfer Transformer (T5)](https://huggingface.co/docs/transformers/en/model_doc/t5) model trained to
|
||||
generate top-k possible queries for which the given document would be an answer.
|
||||
These predicted short queries (up to ~50-60 words) can have repetitions in them,
|
||||
so it also contributes to the frequency of the terms if the term frequency is considered by the retriever.
|
||||
|
||||
The problem with docT5query expansion is a very long inference time, as with any generative model:
|
||||
it can generate only one token per run, and it spends a fair share of resources on it.
|
||||
|
||||
#### External Document Expansion with Term Independent Likelihood MODel (TILDE)
|
||||
|
||||

|
||||
|
||||
[`Term Independent Likelihood MODel (TILDE)`](https://github.com/ielab/TILDE) is an external expansion method that reduces the passage expansion time compared to
|
||||
docT5query by 98%. It uses the assumption that words in texts are independent of each other
|
||||
(as if we were inserting in our speech words without paying attention to their order), which allows for the parallelisation of document expansion.
|
||||
|
||||
Instead of predicting queries, TILDE predicts the most likely terms to see next after reading a passage’s text
|
||||
(**query likelihood paradigm**). TILDE takes the probability distribution of all tokens in a BERT vocabulary based on the document’s text
|
||||
and appends top-k of them to the document without repetitions.
|
||||
|
||||
***Problems of external document expansion:*** External document expansion might not be feasible in many production scenarios where there’s not enough time or compute to expand each and every
|
||||
document you want to store in a database and then additionally do all the calculations needed for retrievers.
|
||||
To solve this problem, a generation of models was developed which do everything in one go, expanding documents “internally”.
|
||||
|
||||
### Internal Document Expansion
|
||||
|
||||
Let’s assume we don’t care about the context of query terms, so we can treat them as independent words that we combine in random order to get
|
||||
the result. Then, for each contextualized term in a document, we are free to pre-compute how this term affects every word in our vocabulary.
|
||||
|
||||
For each document, a vector of the vocabulary length is created. To fill this vector in, for each word in the vocabulary, it is checked if the
|
||||
influence of any document term on it is big enough to consider it. Otherwise, the vocabulary word’s score in a document vector will be zero.
|
||||
For example, by pre-computing vectors for the document “*pizza Margherita*” on a vocabulary of 50,000 most used English words,
|
||||
for this small document of two words, we will get a 50,000-dimensional vector of zeros, where non-zero values will be for a “*pizza*”, “*pizzeria*”,
|
||||
“*flower*”, “*woman*”, “*girl*”, "*Margherita*", “*cocktail*” and “*pizzaiolo*”.
|
||||
|
||||
### Sparse Neural Retriever with Internal Document Expansion
|
||||

|
||||
|
||||
The authors of the [`Sparse Transformer Matching (SPARTA)`](https://arxiv.org/pdf/2009.13013) model use BERT’s model and BERT’s vocabulary (around 30,000 tokens).
|
||||
For each token in BERT vocabulary, they find the maximum dot product between it and contextualized tokens in a document
|
||||
and learn a threshold of a considerable (non-zero) effect.
|
||||
Then, at the inference time, the only thing to be done is to sum up all scores of query tokens in that document.
|
||||
|
||||
***Why is SPARTA not a perfect solution?*** Trained on the MS MARCO dataset, many sparse neural retrievers, including SPARTA,
|
||||
show good results on MS MARCO test data, but when it comes to generalisation (working with other data), they
|
||||
[could perform worse than BM25](https://arxiv.org/pdf/2307.10488).
|
||||
|
||||
### State-of-the-Art of Modern Sparse Neural Retrieval
|
||||
|
||||

|
||||
The authors of the [`Sparse Lexical and Expansion Model (SPLADE)]`](https://arxiv.org/pdf/2109.10086) family of models added dense model training tricks to the
|
||||
internal document expansion idea, which made the retrieval quality noticeably better.
|
||||
|
||||
- The SPARTA model is not sparse enough by construction, so authors of the SPLADE family of models introduced explicit **sparsity regularisation**,
|
||||
preventing the model from producing too many non-zero values.
|
||||
- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specifity of Information Retrieval problem,
|
||||
so SPLADE models introduce a trainable neural network on top of BERT with a specific architecture choice to make it perfectly fit the task.
|
||||
- SPLADE family of models, finally, uses **knowledge distillation**, which is learning from a bigger
|
||||
(and therefore much slower, not-so-fit for production tasks) model how to predict good representations.
|
||||
|
||||
One of the last versions of the SPLADE family of models is [`SPLADE++`](https://arxiv.org/pdf/2205.04733). \
|
||||
SPLADE++, opposed to SPARTA model, expands not only documents but also queries at inference time.
|
||||
We’ll demonstrate this in the next section.
|
||||
|
||||
## SPLADE++ in Qdrant
|
||||
In Qdrant, you can use [`SPLADE++`](https://arxiv.org/pdf/2205.04733) easily with our lightweight library for embeddings called [FastEmbed](https://qdrant.tech/documentation/fastembed/).
|
||||
#### Setup
|
||||
Install `FastEmbed`.
|
||||
|
||||
```python
|
||||
pip install fastembed
|
||||
```
|
||||
|
||||
Import sparse text embedding models supported in FastEmbed.
|
||||
|
||||
```python
|
||||
from fastembed import SparseTextEmbedding
|
||||
```
|
||||
|
||||
You can list all sparse text embedding models currently supported.
|
||||
|
||||
```python
|
||||
SparseTextEmbedding.list_supported_models()
|
||||
```
|
||||
<details>
|
||||
<summary>Output with a list of supported models</summary>
|
||||
|
||||
```bash
|
||||
[{'model': 'prithivida/Splade_PP_en_v1',
|
||||
'vocab_size': 30522,
|
||||
'description': 'Independent Implementation of SPLADE++ Model for English',
|
||||
'size_in_GB': 0.532,
|
||||
'sources': {'hf': 'Qdrant/SPLADE_PP_en_v1'},
|
||||
'model_file': 'model.onnx'},
|
||||
{'model': 'prithvida/Splade_PP_en_v1',
|
||||
'vocab_size': 30522,
|
||||
'description': 'Independent Implementation of SPLADE++ Model for English',
|
||||
'size_in_GB': 0.532,
|
||||
'sources': {'hf': 'Qdrant/SPLADE_PP_en_v1'},
|
||||
'model_file': 'model.onnx'},
|
||||
{'model': 'Qdrant/bm42-all-minilm-l6-v2-attentions',
|
||||
'vocab_size': 30522,
|
||||
'description': 'Light sparse embedding model, which assigns an importance score to each token in the text',
|
||||
'size_in_GB': 0.09,
|
||||
'sources': {'hf': 'Qdrant/all_miniLM_L6_v2_with_attentions'},
|
||||
'model_file': 'model.onnx',
|
||||
'additional_files': ['stopwords.txt'],
|
||||
'requires_idf': True},
|
||||
{'model': 'Qdrant/bm25',
|
||||
'description': 'BM25 as sparse embeddings meant to be used with Qdrant',
|
||||
'size_in_GB': 0.01,
|
||||
'sources': {'hf': 'Qdrant/bm25'},
|
||||
'model_file': 'mock.file',
|
||||
'additional_files': ['arabic.txt',
|
||||
'azerbaijani.txt',
|
||||
'basque.txt',
|
||||
'bengali.txt',
|
||||
'catalan.txt',
|
||||
'chinese.txt',
|
||||
'danish.txt',
|
||||
'dutch.txt',
|
||||
'english.txt',
|
||||
'finnish.txt',
|
||||
'french.txt',
|
||||
'german.txt',
|
||||
'greek.txt',
|
||||
'hebrew.txt',
|
||||
'hinglish.txt',
|
||||
'hungarian.txt',
|
||||
'indonesian.txt',
|
||||
'italian.txt',
|
||||
'kazakh.txt',
|
||||
'nepali.txt',
|
||||
'norwegian.txt',
|
||||
'portuguese.txt',
|
||||
'romanian.txt',
|
||||
'russian.txt',
|
||||
'slovene.txt',
|
||||
'spanish.txt',
|
||||
'swedish.txt',
|
||||
'tajik.txt',
|
||||
'turkish.txt'],
|
||||
'requires_idf': True}]
|
||||
```
|
||||
</details>
|
||||
|
||||
Load SPLADE++.
|
||||
```python
|
||||
sparse_model_name = "prithivida/Splade_PP_en_v1"
|
||||
sparse_model = SparseTextEmbedding(model_name=sparse_model_name)
|
||||
```
|
||||
The model files will be fetched and downloaded, with progress showing.
|
||||
|
||||
#### Embed data
|
||||
We will use a toy movie description dataset.
|
||||
|
||||
<details>
|
||||
<summary> Movie description dataset </summary>
|
||||
|
||||
```python
|
||||
descriptions = ["In 1431, Jeanne d'Arc is placed on trial on charges of heresy. The ecclesiastical jurists attempt to force Jeanne to recant her claims of holy visions.",
|
||||
"A film projectionist longs to be a detective, and puts his meagre skills to work when he is framed by a rival for stealing his girlfriend's father's pocketwatch.",
|
||||
"A group of high-end professional thieves start to feel the heat from the LAPD when they unknowingly leave a clue at their latest heist.",
|
||||
"A petty thief with an utter resemblance to a samurai warlord is hired as the lord's double. When the warlord later dies the thief is forced to take up arms in his place.",
|
||||
"A young boy named Kubo must locate a magical suit of armour worn by his late father in order to defeat a vengeful spirit from the past.",
|
||||
"A biopic detailing the 2 decades that Punjabi Sikh revolutionary Udham Singh spent planning the assassination of the man responsible for the Jallianwala Bagh massacre.",
|
||||
"When a machine that allows therapists to enter their patients' dreams is stolen, all hell breaks loose. Only a young female therapist, Paprika, can stop it.",
|
||||
"An ordinary word processor has the worst night of his life after he agrees to visit a girl in Soho whom he met that evening at a coffee shop.",
|
||||
"A story that revolves around drug abuse in the affluent north Indian State of Punjab and how the youth there have succumbed to it en-masse resulting in a socio-economic decline.",
|
||||
"A world-weary political journalist picks up the story of a woman's search for her son, who was taken away from her decades ago after she became pregnant and was forced to live in a convent.",
|
||||
"Concurrent theatrical ending of the TV series Neon Genesis Evangelion (1995).",
|
||||
"During World War II, a rebellious U.S. Army Major is assigned a dozen convicted murderers to train and lead them into a mass assassination mission of German officers.",
|
||||
"The toys are mistakenly delivered to a day-care center instead of the attic right before Andy leaves for college, and it's up to Woody to convince the other toys that they weren't abandoned and to return home.",
|
||||
"A soldier fighting aliens gets to relive the same day over and over again, the day restarting every time he dies.",
|
||||
"After two male musicians witness a mob hit, they flee the state in an all-female band disguised as women, but further complications set in.",
|
||||
"Exiled into the dangerous forest by her wicked stepmother, a princess is rescued by seven dwarf miners who make her part of their household.",
|
||||
"A renegade reporter trailing a young runaway heiress for a big story joins her on a bus heading from Florida to New York, and they end up stuck with each other when the bus leaves them behind at one of the stops.",
|
||||
"Story of 40-man Turkish task force who must defend a relay station.",
|
||||
"Spinal Tap, one of England's loudest bands, is chronicled by film director Marty DiBergi on what proves to be a fateful tour.",
|
||||
"Oskar, an overlooked and bullied boy, finds love and revenge through Eli, a beautiful but peculiar girl."]
|
||||
```
|
||||
</details>
|
||||
|
||||
Embed movie descriptions with SPLADE++.
|
||||
|
||||
```python
|
||||
sparse_descriptions = list(sparse_model.embed(descriptions))
|
||||
```
|
||||
You can check how a sparse vector generated by SPLADE++ looks in Qdrant.
|
||||
|
||||
```python
|
||||
sparse_descriptions[0]
|
||||
```
|
||||
|
||||
It is stored as **indices** of BERT tokens, weights of which are non-zero, and **values** of these weights.
|
||||
|
||||
```bash
|
||||
SparseEmbedding(
|
||||
values=array([1.57449973, 0.90787691, ..., 1.21796167, 1.1321187]),
|
||||
indices=array([ 1040, 2001, ..., 28667, 29137])
|
||||
)
|
||||
```
|
||||
#### Upload Embeddings to Qdrant
|
||||
Install `qdrant-client`
|
||||
|
||||
```python
|
||||
pip install qdrant-client
|
||||
```
|
||||
|
||||
Qdrant Client has a simple in-memory mode that allows you to experiment locally on small data volumes.
|
||||
Alternatively, you could use for experiments [a free tier cluster](https://qdrant.tech/documentation/cloud/create-cluster/#create-a-cluster)
|
||||
in Qdrant Cloud.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
qdrant_client = QdrantClient(":memory:") # Qdrant is running from RAM.
|
||||
```
|
||||
|
||||
Now, let's create a [collection](https://qdrant.tech/documentation/concepts/collections/) in which could upload our sparse SPLADE++ embeddings. \
|
||||
For that, we will use the [sparse vectors](https://qdrant.tech/documentation/concepts/vectors/#sparse-vectors) representation supported in Qdrant.
|
||||
|
||||
```python
|
||||
qdrant_client.create_collection(
|
||||
collection_name="movies",
|
||||
vectors_config={},
|
||||
sparse_vectors_config={
|
||||
"film_description": models.SparseVectorParams(),
|
||||
},
|
||||
)
|
||||
```
|
||||
To make this collection human-readable, let's save movie metadata (name, description and movie's length) together with an embeddings.
|
||||
<details>
|
||||
<summary> Movie metadata </summary>
|
||||
|
||||
```python
|
||||
metadata = [{"movie_name": "The Passion of Joan of Arc", "movie_watch_time_min": 114, "movie_description": "In 1431, Jeanne d'Arc is placed on trial on charges of heresy. The ecclesiastical jurists attempt to force Jeanne to recant her claims of holy visions."},
|
||||
{"movie_name": "Sherlock Jr.", "movie_watch_time_min": 45, "movie_description": "A film projectionist longs to be a detective, and puts his meagre skills to work when he is framed by a rival for stealing his girlfriend's father's pocketwatch."},
|
||||
{"movie_name": "Heat", "movie_watch_time_min": 170, "movie_description": "A group of high-end professional thieves start to feel the heat from the LAPD when they unknowingly leave a clue at their latest heist."},
|
||||
{"movie_name": "Kagemusha", "movie_watch_time_min": 162, "movie_description": "A petty thief with an utter resemblance to a samurai warlord is hired as the lord's double. When the warlord later dies the thief is forced to take up arms in his place."},
|
||||
{"movie_name": "Kubo and the Two Strings", "movie_watch_time_min": 101, "movie_description": "A young boy named Kubo must locate a magical suit of armour worn by his late father in order to defeat a vengeful spirit from the past."},
|
||||
{"movie_name": "Sardar Udham", "movie_watch_time_min": 164, "movie_description": "A biopic detailing the 2 decades that Punjabi Sikh revolutionary Udham Singh spent planning the assassination of the man responsible for the Jallianwala Bagh massacre."},
|
||||
{"movie_name": "Paprika", "movie_watch_time_min": 90, "movie_description": "When a machine that allows therapists to enter their patients' dreams is stolen, all hell breaks loose. Only a young female therapist, Paprika, can stop it."},
|
||||
{"movie_name": "After Hours", "movie_watch_time_min": 97, "movie_description": "An ordinary word processor has the worst night of his life after he agrees to visit a girl in Soho whom he met that evening at a coffee shop."},
|
||||
{"movie_name": "Udta Punjab", "movie_watch_time_min": 148, "movie_description": "A story that revolves around drug abuse in the affluent north Indian State of Punjab and how the youth there have succumbed to it en-masse resulting in a socio-economic decline."},
|
||||
{"movie_name": "Philomena", "movie_watch_time_min": 98, "movie_description": "A world-weary political journalist picks up the story of a woman's search for her son, who was taken away from her decades ago after she became pregnant and was forced to live in a convent."},
|
||||
{"movie_name": "Neon Genesis Evangelion: The End of Evangelion", "movie_watch_time_min": 87, "movie_description": "Concurrent theatrical ending of the TV series Neon Genesis Evangelion (1995)."},
|
||||
{"movie_name": "The Dirty Dozen", "movie_watch_time_min": 150, "movie_description": "During World War II, a rebellious U.S. Army Major is assigned a dozen convicted murderers to train and lead them into a mass assassination mission of German officers."},
|
||||
{"movie_name": "Toy Story 3", "movie_watch_time_min": 103, "movie_description": "The toys are mistakenly delivered to a day-care center instead of the attic right before Andy leaves for college, and it's up to Woody to convince the other toys that they weren't abandoned and to return home."},
|
||||
{"movie_name": "Edge of Tomorrow", "movie_watch_time_min": 113, "movie_description": "A soldier fighting aliens gets to relive the same day over and over again, the day restarting every time he dies."},
|
||||
{"movie_name": "Some Like It Hot", "movie_watch_time_min": 121, "movie_description": "After two male musicians witness a mob hit, they flee the state in an all-female band disguised as women, but further complications set in."},
|
||||
{"movie_name": "Snow White and the Seven Dwarfs", "movie_watch_time_min": 83, "movie_description": "Exiled into the dangerous forest by her wicked stepmother, a princess is rescued by seven dwarf miners who make her part of their household."},
|
||||
{"movie_name": "It Happened One Night", "movie_watch_time_min": 105, "movie_description": "A renegade reporter trailing a young runaway heiress for a big story joins her on a bus heading from Florida to New York, and they end up stuck with each other when the bus leaves them behind at one of the stops."},
|
||||
{"movie_name": "Nefes: Vatan Sagolsun", "movie_watch_time_min": 128, "movie_description": "Story of 40-man Turkish task force who must defend a relay station."},
|
||||
{"movie_name": "This Is Spinal Tap", "movie_watch_time_min": 82, "movie_description": "Spinal Tap, one of England's loudest bands, is chronicled by film director Marty DiBergi on what proves to be a fateful tour."},
|
||||
{"movie_name": "Let the Right One In", "movie_watch_time_min": 114, "movie_description": "Oskar, an overlooked and bullied boy, finds love and revenge through Eli, a beautiful but peculiar girl."}]
|
||||
```
|
||||
</details>
|
||||
|
||||
Upload embedded descriptions with movie metadata into the collection.
|
||||
```python
|
||||
qdrant_client.upsert(
|
||||
collection_name="movies",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=idx,
|
||||
payload=metadata[idx],
|
||||
vector={
|
||||
"film_description": models.SparseVector(
|
||||
indices=vector.indices,
|
||||
values=vector.values
|
||||
)
|
||||
},
|
||||
)
|
||||
for idx, vector in enumerate(sparse_descriptions)
|
||||
],
|
||||
)
|
||||
```
|
||||
#### Querying
|
||||
Let’s query our collection!
|
||||
|
||||
```python
|
||||
query_embedding = list(sparse_model.embed("A movie about music"))[0]
|
||||
|
||||
response = qdrant_client.query_points(
|
||||
collection_name="movies",
|
||||
query=models.SparseVector(indices=query_embedding.indices, values=query_embedding.values),
|
||||
using="film_description",
|
||||
limit=1,
|
||||
with_vectors=True,
|
||||
with_payload=True
|
||||
)
|
||||
print(response)
|
||||
```
|
||||
Output looks like this:
|
||||
```bash
|
||||
points=[ScoredPoint(
|
||||
id=18,
|
||||
version=0,
|
||||
score=9.6779785,
|
||||
payload={
|
||||
'movie_name': 'This Is Spinal Tap',
|
||||
'movie_watch_time_min': 82,
|
||||
'movie_description': "Spinal Tap, one of England's loudest bands,
|
||||
is chronicled by film director Marty DiBergi on what proves to be a fateful tour."
|
||||
},
|
||||
vector={
|
||||
'film_description': SparseVector(
|
||||
indices=[1010, 2001, ..., 25316, 25517],
|
||||
values=[0.49717945, 0.19760133, ..., 1.2124698, 0.58689135])
|
||||
},
|
||||
shard_key=None,
|
||||
order_value=None
|
||||
)]
|
||||
```
|
||||
As you can see, there are no overlapping words in the query and a description of a found movie,
|
||||
even though the answer fits the query, and yet we’re working with **exact matching**. \
|
||||
This is possible due to the **internal expansion** of the query and the document that SPLADE++ does.
|
||||
|
||||
#### Internal Expansion by SPLADE++
|
||||
|
||||
Let’s check how did SPLADE++ expand the query and the document we got as an answer. \
|
||||
For that, we will need to use the HuggingFace library called [Tokenizers](https://huggingface.co/docs/tokenizers/en/index).
|
||||
With it, we will be able to decode back to human-readable format **indices** of words in a vocabulary SPLADE++ uses.
|
||||
|
||||
Firstly we will need to install this library.
|
||||
|
||||
```python
|
||||
pip install tokenizers
|
||||
```
|
||||
|
||||
Then, let's write a function which will decode SPLADE++ sparse embeddings and return words SPLADE++ uses for encoding the input. \
|
||||
We would like to return them in the descending order based on the weight (**impact score**), SPLADE++ assigned them.
|
||||
|
||||
```python
|
||||
from tokenizers import Tokenizer
|
||||
|
||||
tokenizer = Tokenizer.from_pretrained('Qdrant/SPLADE_PP_en_v1')
|
||||
|
||||
def get_tokens_and_weights(sparse_embedding, tokenizer):
|
||||
token_weight_dict = {}
|
||||
for i in range(len(sparse_embedding.indices)):
|
||||
token = tokenizer.decode([sparse_embedding.indices[i]])
|
||||
weight = sparse_embedding.values[i]
|
||||
token_weight_dict[token] = weight
|
||||
|
||||
# Sort the dictionary by weights
|
||||
token_weight_dict = dict(sorted(token_weight_dict.items(), key=lambda item: item[1], reverse=True))
|
||||
return token_weight_dict
|
||||
```
|
||||
|
||||
Firstly, we apply our function to the query.
|
||||
|
||||
```python
|
||||
query_embedding = list(sparse_model.embed("A movie about music"))[0]
|
||||
print(get_tokens_and_weights(query_embedding, tokenizer))
|
||||
```
|
||||
That’s how SPLADE++ expanded the query:
|
||||
```bash
|
||||
{
|
||||
"music": 2.764289617538452,
|
||||
"movie": 2.674748420715332,
|
||||
"film": 2.3489091396331787,
|
||||
"musical": 2.276120901107788,
|
||||
"about": 2.124547004699707,
|
||||
"movies": 1.3825485706329346,
|
||||
"song": 1.2893378734588623,
|
||||
"genre": 0.9066758751869202,
|
||||
"songs": 0.8926399946212769,
|
||||
"a": 0.8900706768035889,
|
||||
"musicians": 0.5638002157211304,
|
||||
"sound": 0.49310919642448425,
|
||||
"musician": 0.46415239572525024,
|
||||
"drama": 0.462990403175354,
|
||||
"tv": 0.4398191571235657,
|
||||
"book": 0.38950803875923157,
|
||||
"documentary": 0.3758136034011841,
|
||||
"hollywood": 0.29099565744400024,
|
||||
"story": 0.2697228491306305,
|
||||
"nature": 0.25306591391563416,
|
||||
"concerning": 0.205053448677063,
|
||||
"game": 0.1546829640865326,
|
||||
"rock": 0.11775632947683334,
|
||||
"definition": 0.08842901140451431,
|
||||
"love": 0.08636035025119781,
|
||||
"soundtrack": 0.06807517260313034,
|
||||
"religion": 0.053535860031843185,
|
||||
"filmed": 0.025964470580220222,
|
||||
"sounds": 0.0004048719711136073
|
||||
}
|
||||
```
|
||||
|
||||
Then, we apply our function to the answer.
|
||||
```python
|
||||
query_embedding = list(sparse_model.embed("A movie about music"))[0]
|
||||
|
||||
response = qdrant_client.query_points(
|
||||
collection_name="movies",
|
||||
query=models.SparseVector(indices=query_embedding.indices, values=query_embedding.values),
|
||||
using="film_description",
|
||||
limit=1,
|
||||
with_vectors=True,
|
||||
with_payload=True
|
||||
)
|
||||
|
||||
print(get_tokens_and_weights(response.points[0].vector['film_description'], tokenizer))
|
||||
```
|
||||
|
||||
And that's how SPLADE++ expanded the answer.
|
||||
|
||||
```python
|
||||
{'spinal': 2.6548674, 'tap': 2.534881, 'marty': 2.223297, '##berg': 2.0402722,
|
||||
'##ful': 2.0030282, 'fate': 1.935915, 'loud': 1.8381964, 'spine': 1.7507898,
|
||||
'di': 1.6161551, 'bands': 1.5897619, 'band': 1.589473, 'uk': 1.5385966, 'tour': 1.4758654,
|
||||
'chronicle': 1.4577943, 'director': 1.4423795, 'england': 1.4301306, '##est': 1.3025658,
|
||||
'taps': 1.2124698, 'film': 1.1069428, '##berger': 1.1044296, 'tapping': 1.0424755, 'best': 1.0327196,
|
||||
'louder': 0.9229055, 'music': 0.9056678, 'directors': 0.8887502, 'movie': 0.870712, 'directing': 0.8396196,
|
||||
'sound': 0.83609974, 'genre': 0.803052, 'dave': 0.80212915, 'wrote': 0.7849579, 'hottest': 0.7594193, 'filmed': 0.750105,
|
||||
'english': 0.72807616, 'who': 0.69502294, 'tours': 0.6833075, 'club': 0.6375339, 'vertebrae': 0.58689135, 'chronicles': 0.57296354,
|
||||
'dance': 0.57278687, 'song': 0.50987065, ',': 0.49717945, 'british': 0.4971719, 'writer': 0.495709, 'directed': 0.4875775,
|
||||
'cork': 0.475757, '##i': 0.47122696, '##band': 0.46837863, 'most': 0.44112885, '##liest': 0.44084555, 'destiny': 0.4264851,
|
||||
'prove': 0.41789067, 'is': 0.40306947, 'famous': 0.40230379, 'hop': 0.3897451, 'noise': 0.38770816, '##iest': 0.3737782,
|
||||
'comedy': 0.36903998, 'sport': 0.35883865, 'quiet': 0.3552795, 'detail': 0.3397654, 'fastest': 0.30345848, 'filmmaker': 0.3013101,
|
||||
'festival': 0.28146765, '##st': 0.28040633, 'tram': 0.27373192, 'well': 0.2599603, 'documentary': 0.24368097, 'beat': 0.22953634,
|
||||
'direction': 0.22925079, 'hardest': 0.22293334, 'strongest': 0.2018861, 'was': 0.19760133, 'oldest': 0.19532987,
|
||||
'byron': 0.19360808, 'worst': 0.18397793, 'touring': 0.17598206, 'rock': 0.17319143, 'clubs': 0.16090117,
|
||||
'popular': 0.15969758, 'toured': 0.15917331, 'trick': 0.1530599, 'celebrity': 0.14458777, 'musical': 0.13888633,
|
||||
'filming': 0.1363699, 'culture': 0.13616633, 'groups': 0.1340591, 'ski': 0.13049376, 'venue': 0.12992987,
|
||||
'style': 0.12853126, 'history': 0.12696269, 'massage': 0.11969914, 'theatre': 0.11673525, 'sounds': 0.108338095,
|
||||
'visit': 0.10516077, 'editing': 0.078659914, 'death': 0.066746496, 'massachusetts': 0.055702563, 'stuart': 0.0447934,
|
||||
'romantic': 0.041140396, 'pamela': 0.03561337, 'what': 0.016409796, 'smallest': 0.010815808, 'orchestra': 0.0020691194}
|
||||
```
|
||||
Due to the expansion both the query and the document overlap in “*music*”, “*film*”, “*sounds*”,
|
||||
and others, so **exact matching** works.
|
||||
|
||||
## Key Takeaways: When to Choose Sparse Neural Models for Retrieval
|
||||
Sparse Neural Retrieval makes sense:
|
||||
|
||||
- In areas where keyword matching is crucial but BM25 is insufficient for initial retrieval, semantic matching (e.g., synonyms, homonyms) adds significant value. This is especially true in fields such as medicine, academia, law, and e-commerce, where brand names and serial numbers play a critical role. Dense retrievers tend to return many false positives, while sparse neural retrieval helps narrow down these false positives.
|
||||
|
||||
- Sparse neural retrieval can be a valuable option for scaling, especially when working with large datasets. It leverages exact matching using an inverted index, which can be fast depending on the nature of your data.
|
||||
|
||||
- If you’re using traditional retrieval systems, sparse neural retrieval is compatible with them and helps bridge the semantic gap.
|
||||
|
||||
@@ -347,7 +347,7 @@ query_vec, query_tokens = compute_vector(query_text)
|
||||
query_vec.shape
|
||||
|
||||
query_indices = query_vec.nonzero().numpy().flatten()
|
||||
query_values = query_vec.detach().numpy()[indices]
|
||||
query_values = query_vec.detach().numpy()[query_indices]
|
||||
```
|
||||
|
||||
In this example, we use the same model for both document and query. This is not a requirement, but it's a simpler approach.
|
||||
|
||||
@@ -153,7 +153,7 @@ client.create_collection(
|
||||
)
|
||||
```
|
||||
|
||||
For other configurations like `hnsw_config.on_disk` or `memmap_threshold_kb`, see the Qdrant documentation for [Storage.](https://qdrant.tech/documentation/concepts/storage/)
|
||||
For other configurations like `hnsw_config.on_disk` or `memmap_threshold`, see the Qdrant documentation for [Storage.](https://qdrant.tech/documentation/concepts/storage/)
|
||||
|
||||
### SDKs
|
||||
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
---
|
||||
draft: false
|
||||
title: "Empowering QA.tech’s Testing Agents with Real-Time Precision and Scale"
|
||||
short_description: "Using vectors to enable scalable & real-time QA testing."
|
||||
description: "QA.tech uses Qdrant to power AI agents, enabling scalable, real-time web testing with custom embeddings and batch efficiency."
|
||||
preview_image: /blog/case-study-qatech/preview.png
|
||||
social_preview_image: /blog/case-study-qatech/preview.png
|
||||
date: 2024-11-21T00:02:00Z
|
||||
author: Qdrant
|
||||
featured: false
|
||||
tags:
|
||||
- Cuality Assurance
|
||||
- Vector Search
|
||||
- Automated Testing
|
||||
- PGVector
|
||||
- Agents
|
||||
- Retrieval Augmented Generation
|
||||
---
|
||||
|
||||

|
||||
|
||||
[QA.tech](https://qa.tech/), a company specializing in AI-driven automated testing solutions, found that building and **fully testing web applications, especially end-to-end, can be complex and time-consuming**. Unlike unit tests, end-to-end tests reveal what’s actually happening in the browser, often uncovering issues that other methods miss.
|
||||
|
||||
Traditional solutions like hard-coded tests are not only labor-intensive to set up but also challenging to maintain over time. Alternatively, hiring QA testers can be a solution, but for startups, it quickly becomes a bottleneck. With every release, more testers are needed, and if testing is outsourced, managing timelines and ensuring quality becomes even harder.
|
||||
|
||||
To address this, QA.tech has developed **testing agents** that perform tasks on the browser just like a user would - for example, purchasing a ticket on a travel app. These agents navigate the entire booking process, from searching for flights to completing the purchase, all while assessing their success. **They document errors, record the process, and flag issues for developers to review.** With access to console logs and network calls, developers can easily analyze each step, quickly understanding and debugging any issues that arise.
|
||||
|
||||

|
||||
|
||||
*Output from a QA.tech AI agent*
|
||||
|
||||
## What prompted QA.tech to use a vector database?
|
||||
|
||||
QA.tech initially used **pgvector** for simpler vector use cases but encountered scalability limitations as their requirements grew, prompting them to adopt Qdrant. They needed a [vector database](/qdrant-vector-database/) capable of handling high-velocity, real-time analysis to support their AI agents, which operate within an analysis layer that observes and interprets actions across web pages. This analysis layer relies heavily on multimodal models and substantial subprocessing to enable the AI agent to make informed, real-time decisions.
|
||||
|
||||
In some web interfaces, hundreds of actions can occur, and processing them in real time - especially with each click - can be slow. Dynamic web elements and changing identifiers further complicate this, making traditional methods unreliable. To address these challenges, QA.tech trained custom embeddings on specific actions, which significantly accelerates decision-making.
|
||||
|
||||
This setup requires frequent embedding lookups, generating a high volume of database calls for each interaction. As **Vilhelm von Ehrenheim from QA.tech** explained:
|
||||
|
||||
> “You get a lot of embeddings, a lot of calls, a lot of lookups towards the database for every click, and that needs to scale nicely.”
|
||||
|
||||
Qdrant’s fast, scalable [vector search](/advanced-search/) enables QA.tech to handle these high-velocity lookups seamlessly, ensuring that the agent remains responsive and capable of making quick, accurate decisions in real time.
|
||||
|
||||
## Why QA.tech chose Qdrant for its AI Agent platform
|
||||
|
||||
QA.tech’s AI Agents handle high-velocity web actions, requiring efficient real-time operations and scalable infrastructure. The team faced challenges with managing network overhead, CPU load, and the need to store [multiple embeddings](/documentation/concepts/vectors/#multivectors) for different use cases. Qdrant provided the solution to address these issues.
|
||||
|
||||
**Reducing Network Overhead with Batch Operations**
|
||||
|
||||
Handling hundreds of simultaneous actions on a web interface individually created significant network overhead. Von Ehrenheim explained that “doing all of those in separate calls creates a lot of network overhead.” Qdrant’s batch operations allowed QA.tech to process multiple actions at once, reducing network traffic and improving efficiency. This capability is essential for AI Agents, where real-time responsiveness is critical.
|
||||
|
||||
**Optimizing CPU Load for Embedding Processing**
|
||||
|
||||
PostgreSQL’s transaction guarantees resulted in high CPU usage when processing embeddings, especially at scale. Von Ehrenheim noted that adding many new embeddings "requires much more CPU," which led to performance bottlenecks. Qdrant’s architecture efficiently handled large-scale embeddings, preventing CPU overload and ensuring smooth, uninterrupted performance, a key requirement for AI Agents.
|
||||
|
||||
**Managing Multiple Embeddings for Different Use Cases**
|
||||
|
||||
AI Agents need flexibility in handling both real-time actions and context-aware tasks. QA.tech required different embeddings for immediate action processing and deeper semantic searches. Von Ehrenheim mentioned, *“We use one embedding for high-velocity actions, but I also want to store other types of embeddings for analytical purposes.”*
|
||||
|
||||
> Qdrant’s ability to store multiple embeddings per data point allowed QA.tech to meet these diverse needs without added complexity.
|
||||
|
||||
|
||||
## How QA.tech Overcame Key Challenges in AI Agent Development
|
||||
|
||||
Building reliable AI agents presents unique complexities, particularly as workflows grow more multi-step and dynamic.
|
||||
|
||||
> "The more steps you ask an agent to take, the harder it becomes to ensure consistent performance," Vilhelm von Ehrenheim, Co-Founder of QA.tech.
|
||||
|
||||
Each additional action adds layers of interdependent variables, creating pathways that can easily lead to errors if not managed carefully.
|
||||
|
||||
Von Ehrenheim also points out the limitations of current large language models (LLMs), noting that *“LLMs are getting more powerful, but they still struggle with multi-step reasoning and for example handling subtle visual changes like dark mode or adaptive UIs.”* These challenges make it essential for agents to have precise planning capabilities and context awareness, which QA.tech has addressed by implementing custom embeddings and multimodal models.
|
||||
|
||||
*“This is where scalable, adaptable infrastructure becomes crucial,”* von Ehrenheim adds. Qdrant has been instrumental for QA.tech, providing stable, high-performance vector search to support the demanding workflows. **“With Qdrant, we’re able to handle these complex, high-velocity tasks without compromising on reliability.”**
|
||||
@@ -0,0 +1,149 @@
|
||||
---
|
||||
draft: false
|
||||
title: "How Sprinklr Leverages Qdrant to Enhance AI-Driven Customer Experience Solutions"
|
||||
short_description: "Using vector search to power AI-driven tools for customer engagement."
|
||||
description: "Learn how Sprinklr uses vector search to power AI tools for customer engagement."
|
||||
preview_image: /blog/case-study-sprinklr/preview.png
|
||||
social_preview_image: /blog/case-study-sprinklr/preview.png
|
||||
date: 2024-10-17T00:02:00Z
|
||||
author: Qdrant
|
||||
featured: false
|
||||
tags:
|
||||
- Sprinklr
|
||||
- Qdrant
|
||||
- AI Search
|
||||
- Qdrant Benchmarks
|
||||
- ElasticSearch
|
||||
- Vector Search
|
||||
---
|
||||
|
||||

|
||||
|
||||
|
||||
[Sprinklr](https://www.sprinklr.com/), a leader in unified customer experience management (Unified-CXM), helps global brands engage customers meaningfully across more than 30 digital channels. To achieve this, Sprinklr needed a scalable solution for AI-powered search to support their AI applications, particularly in handling the vast data requirements of customer interactions.
|
||||
|
||||
Raghav Sonavane, Associate Director of Machine Learning Engineering at Sprinklr, leads the Applied AI team, focusing on Generative AI (GenAI) and Retrieval-Augmented Generation (RAG). His team is responsible for training and fine-tuning in-house models and deploying advanced retrieval and generation systems for customer-facing applications like FAQ bots and other [GenAI-driven services](https://www.sprinklr.com/blog/how-sprinklr-uses-RAG/). The team provides all of these capabilities in a centralized platform to the Sprinklr product engineering teams.
|
||||
|
||||

|
||||
|
||||
*Figure:* Sprinklr’s RAG architecture
|
||||
|
||||
Sprinklr’s platform is composed of four key product suites - Sprinklr Service, Sprinklr Marketing, Sprinklr Social, and Sprinklr Insights. Each suite is embedded with AI-first features such as assist agents, post-call analysis, and real-time analytics, which are crucial for managing large-scale contact center operations. “These AI-driven capabilities, supported by Qdrant’s advanced vector search, enhance Sprinklr’s customer-facing tools such as FAQ bots, transactional bots, conversational services, and product recommendation engines,” says Sonavane.
|
||||
|
||||
These self-serve applications rely heavily on advanced vector search to analyze and optimize community content and refine knowledge bases, ensuring efficient and relevant responses. For customers requiring further assistance, Sprinklr equips support agents with powerful search capabilities, enabling them to quickly access similar cases and draw from past interactions, enhancing the quality and speed of customer support.
|
||||
|
||||
## The Need for a Vector Database
|
||||
|
||||
To support various AI-driven applications, Sprinklr needed an efficient vector database. "The key challenge was to provide the highest quality and fastest search capabilities for retrieval tasks across the board," explains Sonavane.
|
||||
|
||||
Last year, Sprinklr undertook a comprehensive evaluation of its existing search infrastructure. The goals were to identify current capability gaps, benchmark performance for speed and cost, and explore opportunities to improve the developer experience through enhanced scalability and stronger data privacy controls. It became clear that an advanced vector database was essential to meet these needs, and Qdrant emerged as the ideal solution.
|
||||
|
||||
### Why Qdrant?
|
||||
|
||||
After evaluating several options of vector DBs, including Pinecone, Weaviate, and ElasticSearch, Sprinklr chose Qdrant for its:
|
||||
|
||||
- **Developer-Friendly Documentation:** “Qdrant’s clear [documentation](https://qdrant.tech/documentation/) enabled our team to integrate it quickly into our workflows,” notes Sonavane.
|
||||
- **High Customizability:** Qdrant provided Sprinklr with essential flexibility through high-level abstractions that allowed for extensive customizations. The diverse teams at Sprinklr, working on various GenAI applications, needed a solution that could adapt to different workloads. “The ability to fine-tune configurations at the collection level was crucial for our varied AI applications,” says Sonavane. Qdrant met this need by offering:
|
||||
|
||||
- **Configuration for high-speed search** that fine-tunes settings for optimal performance.
|
||||
- [**Quantized vectors**](https://qdrant.tech/documentation/guides/quantization/) for high-dimensional data workloads
|
||||
- [**Memory map**](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) for efficient search optimizing memory usage.
|
||||
- **Speed and Cost Efficiency:** Qdrant provided the best combination of speed and cost, making it the most viable solution for Sprinklr’s needs. “We needed a solution that wouldn’t just meet our performance requirements but also keep costs in check, and Qdrant delivered on both fronts,” says Sonavane.
|
||||
- **Enhanced Monitoring:** Qdrant’s monitoring tools further boosted system efficiency, allowing Sprinklr to maintain high performance across their platforms.
|
||||
|
||||
## Implementation and Qdrant’s Performance
|
||||
|
||||
Sprinklr’s transition to Qdrant was carefully managed, starting with 10% of their workloads before gradually scaling up. The transition was seamless, thanks in part to Qdrant’s configurable [Web UI](https://qdrant.tech/documentation/interfaces/web-ui/), which allowed Sprinklr to fully utilize its capabilities within the existing infrastructure.
|
||||
|
||||
“Qdrant’s ability to index [multiple vectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors) simultaneously and retrieve and re-rank with precision brought significant improvements to our workflow,” Sonavane remarks. This feature reduced the need for repeated retrieval processes, significantly improving efficiency. Additionally, Qdrant’s [quantization](https://qdrant.tech/documentation/guides/quantization/) and [memory mapping](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) features enabled Sprinklr to reduce RAM usage, leading to substantial cost savings.
|
||||
|
||||
Qdrant now plays a key supportive role in enhancing Sprinklr’s vector search capabilities within its AI-driven applications, which is designed to be cloud- and LLM-agnostic. The platform supports various AI-driven tasks, from retrieval and re-ranking to serving advanced customer experiences. “Retrieval is the foundation of all our AI tasks, and Qdrant’s resilience and speed have made it an integral part of our system,” Sonavane emphasizes. Sprinklr operates [Qdrant as a managed service on AWS](https://qdrant.tech/cloud/), ensuring scalability, reliability, and ease of use.
|
||||
|
||||
### Key Outcomes with Qdrant
|
||||
|
||||
After rigorous internal evaluation, Sprinklr achieved the following results with Qdrant:
|
||||
|
||||
- **30% Cost Reduction**: Internal benchmarking showed Qdrant reduced Sprinklr's retrieval infrastructure costs by 30%.
|
||||
- **Improved Developer Efficiency**: Qdrant’s user-friendly environment made it easier to maintain instances, enhancing overall efficiency.
|
||||
|
||||
The Sprinklr team conducted a thorough internal benchmark on applications requiring vector search across 10k to over 1M vectors with varying dimensions of vectors depending on the use case. The key results from these benchmarks include:
|
||||
|
||||
- **Superior Write Performance**: Qdrant's write performance excelled in Sprinklr’s benchmark tests, with incremental indexing time for 100k to 1M vectors being less than 10% of Elasticsearch’s, making it highly efficient for handling updates and append queries in high-ingestion use cases.
|
||||
- **Low Latency for Real-Time Applications:** In Sprinklr's benchmark, Qdrant delivered a P99 latency of 20ms for searches on 1 million vectors, making it ideal for real-time use cases like live chat, where Elasticsearch and Milvus both exceeded 100ms.
|
||||
- **High Throughput for Heavy Query Loads**: In Sprinklr's benchmark, Qdrant handled up to 250 requests per second (RPS) under similar configurations, significantly outperforming Elasticsearch's 100 RPS, making it ideal for environments with heavy query loads.
|
||||
|
||||
“Qdrant is a very fast and high quality retrieval system,” Sonavane points out.
|
||||
|
||||

|
||||
|
||||
*Figure: P95 Query Time vs Mean Average Precision Benchmark Across Varying Index Sizes*
|
||||
|
||||
## Outlook
|
||||
|
||||
Looking ahead, the Applied AI team at Sprinklr is focused on developing Sprinklr Digital Twin technology for companies, organizations, and employees, aiming to seamlessly integrate AI agents with human workers in business processes. Sprinklr Digital Twins are powered by a process engine that incorporates personas, skills, tasks, and activities, designed to optimize operational efficiency.
|
||||
|
||||

|
||||
|
||||
*Figure: Sprinklr Digital Twin*
|
||||
|
||||
Vector search will play a crucial role, as each AI agent will have its own knowledge base, skill set, and tool set, enabling precise and autonomous task execution. The integration of Qdrant further enhances the system's ability to manage and utilize large volumes of data effectively.
|
||||
|
||||
|
||||
## Benchmarking Conclusion
|
||||
|
||||
***Configuration Details:***
|
||||
|
||||
- We benchmarked applications requiring search on different sizes ranging from 10k to 1M+ vectors, with varying dimensions of vectors depending on the usage. Our infrastructure mainly consisted of Elasticsearch and in-memory Faiss vector search.
|
||||
|
||||
Key Observations:
|
||||
|
||||
1. **Indexing Speed**: Qdrant indexes vectors rapidly, making it suitable for applications that require quick data ingestion. Among the alternatives tried, milvus was on par with qdrant in terms of indexing time for a given precision. The latest versions of Elasticsearch offer much improvement compared to previous versions, though not as efficient as Qdrant.
|
||||
- **Write Performance:** For some of our use cases, update queries and append queries were significantly higher. For ES, an increase in the number of points had a severe impact on total upload time. For 100k to 1M vector index qdrant incremental indexing time was less than 10% of Elasticsearch.
|
||||
2. **Low Latency**: Tail latencies are very critical for real-time applications such as live chat, requiring low P95 and P99 latencies. For a workload requiring search on 1 million vectors, qdrant provided inference latency of 20ms P99 whereas ES and Milvus were more than 100ms.
|
||||
3. **High Throughput**: Qdrant handles a high number of requests per second, making it ideal for environments with heavy query loads. For similar configurations, Qdrant provided a throughput of 250 RPS whereas ES was around 100 RPS.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
```json
|
||||
data = [
|
||||
|
||||
{'system': 'Qdrant', 'index_size': '1,000', 'MAP': 0.98, 'P95 Time': 0.22, 'Mean Time': 0.1, 'QPS': 280,
|
||||
|
||||
'Upload Time': 1},
|
||||
|
||||
{'system': 'Qdrant', 'index_size': '10,000', 'MAP': 0.99, 'P95 Time': 0.16, 'Mean Time': 0.09, 'QPS': 330,
|
||||
|
||||
'Upload Time': 5},
|
||||
|
||||
{'system': 'Qdrant', 'index_size': '100,000', 'MAP': 0.98, 'P95 Time': 0.3, 'Mean Time': 0.23, 'QPS': 145,
|
||||
|
||||
'Upload Time': 100},
|
||||
|
||||
{'system': 'Qdrant', 'index_size': '1,000,000', 'MAP': 0.99, 'P95 Time': 0.171, 'Mean Time': 0.162, 'QPS': 596,
|
||||
|
||||
'Upload Time': 220},
|
||||
|
||||
{'system': 'ElasticSearch', 'index_size': '1,000', 'MAP': 0.99, 'P95 Time': 0.42, 'Mean Time': 0.32, 'QPS': 95,
|
||||
|
||||
'Upload Time': 10},
|
||||
|
||||
{'system': 'ElasticSearch', 'index_size': '10,000', 'MAP': 0.98, 'P95 Time': 0.3, 'Mean Time': 0.24, 'QPS': 120,
|
||||
|
||||
'Upload Time': 50},
|
||||
|
||||
{'system': 'ElasticSearch', 'index_size': '100,000', 'MAP': 0.99, 'P95 Time': 0.48, 'Mean Time': 0.42, 'QPS': 80,
|
||||
|
||||
'Upload Time': 1100},
|
||||
|
||||
{'system': 'ElasticSearch', 'index_size': '1,000,000', 'MAP': 0.99, 'P95 Time': 0.37, 'Mean Time': 0.236,
|
||||
|
||||
'QPS': 348, 'Upload Time': 1150}
|
||||
|
||||
]
|
||||
```
|
||||
@@ -0,0 +1,117 @@
|
||||
---
|
||||
draft: false
|
||||
title: "Advanced Retrieval with ColPali & Qdrant Vector Database"
|
||||
short_description: "Redefining document retrieval with vision language models."
|
||||
description: "ColPali leverages vision language models and multivector embeddings to streamline complex document retrieval."
|
||||
preview_image: /blog/qdrant-colpali/preview.png
|
||||
social_preview_image: /blog/qdrant-colpali/preview.png
|
||||
date: 2024-11-05T00:02:00Z
|
||||
author: Sabrina Aquino
|
||||
featured: true
|
||||
tags:
|
||||
- ColPali
|
||||
- Qdrant
|
||||
- Document Retrieval
|
||||
- Vision Language Models
|
||||
- Binary Quantization
|
||||
---
|
||||
| Time: 30 min | Level: Advanced | Notebook: [GitHub](https://github.com/qdrant/examples/blob/master/colpali-and-binary-quantization/colpali_demo_binary.ipynb) |
|
||||
| --- | ----------- | ----------- |
|
||||
|
||||
It’s no secret that even the most modern document retrieval systems have a hard time handling visually rich documents like **PDFs, containing tables, images, and complex layouts.**
|
||||
|
||||
ColPali introduces a multimodal retrieval approach that uses **Vision Language Models (VLMs)** instead of the traditional OCR and text-based extraction.
|
||||
|
||||
By processing document images directly, it creates **multi-vector embeddings** from both the visual and textual content, capturing the document's structure and context more effectively. This method outperforms traditional techniques, as demonstrated by the [**Visual Document Retrieval Benchmark (ViDoRe)**](https://huggingface.co/vidore).
|
||||
|
||||
**Before we go any deeper, watch our short video:**
|
||||
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/_A90A-grwIc?si=ezEjuiRJtGZ87yd1" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
## Standard Retrieval vs ColPali
|
||||
|
||||
The standard approach starts by running **Optical Character Recognition (OCR)** to extract the text from a document. Once the text is extracted, a layout detection model interprets the structure, which is followed by chunking the text into smaller sections for embedding. This method works adequately for documents where the text content is the primary focus.
|
||||
|
||||
Imagine you have a PDF packed with complex layouts, tables, and images, and you need to extract meaningful information efficiently. Traditionally, this would involve several steps:
|
||||
|
||||
1. **Text Extraction:** Using OCR to pull words from each page.
|
||||
2. **Layout Detection:** Identifying page elements like tables, paragraphs, and titles.
|
||||
3. **Chunking:** Experimenting with methods to determine the best fit for your use case.
|
||||
4. **Embedding Creation:** Finally generating and storing the embeddings.
|
||||
|
||||
### Why is ColPali Better?
|
||||
|
||||
This entire process can require too many steps, especially for complex documents, with each page often taking over seven seconds to process. For text-heavy documents, this approach might suffice, but real-world data is often rich and complex, making traditional extraction methods less effective.
|
||||
|
||||
This is where ColPali comes into play. **ColPali, or Contextualized Late Interaction Over PaliGemma**, uses a vision language model (VLM) to simplify and enhance the document retrieval process.
|
||||
|
||||
Instead of relying on text-only methods, ColPali generates contextualized **multivector embeddings** directly from an image of a document page. The VLM considers visual elements, structure, and text all at once, creating a holistic representation of each page.
|
||||
|
||||
## How ColPali Works Under the Hood
|
||||

|
||||
|
||||
Rather than relying on OCR, ColPali **processes the entire document as an image** using a Vision Encoder. It creates multi-vector embeddings that capture both the textual content and the visual structure of the document which are then passed through a Large Language Model (LLM), which integrates the information into a representation that retains both text and visual features.
|
||||
|
||||
Here’s a step-by-step look at the ColPali architecture and how it enhances document retrieval:
|
||||
|
||||
1. **Image Preprocessing:** The input image is divided into a 32x32 grid, resulting in 1,024 patches.
|
||||
2. **Contextual Transformation:** Each patch undergoes transformations to capture local and global context and is represented by a 128-dimensional vector.
|
||||
3. **Query Processing:** When a text query is sent, ColPali generates token-level embeddings for the query, comparing it with document patches using a similarity matrix (specifically MaxSim).
|
||||
4. **MaxSim Similarity:** This similarity matrix computes similarities for each query token in every document patch, selecting maximum similarities to efficiently retrieve relevant pages. This late interaction approach helps ColPali capture intricate context across a document’s structure and text.
|
||||
|
||||
> ColPali’s late interaction strategy is inspired by ColBERT and improves search by analyzing layout and textual content in a single pass.
|
||||
|
||||
## Optimizing with Binary Quantization
|
||||

|
||||
|
||||
Binary Quantization further enhances the ColPali pipeline by **reducing storage and computational load** without compromising search performance. Binary Quantization, unlike Scalar Quantization, compresses vectors more aggressively, which can speed up search times and reduce memory usage.
|
||||
|
||||
In an experiment based on a [**blog post by Daniel Van Strien**](https://danielvanstrien.xyz/posts/post-with-code/colpali-qdrant/2024-10-02_using_colpali_with_qdrant.html), where ColPali and Qdrant were used to search a UFO document dataset, the results were compelling. By using Binary Quantization along with rescoring and oversampling techniques, we saw search time reduced by nearly half compared to Scalar Quantization, while maintaining similar accuracy.
|
||||
|
||||
## Using ColPali with Qdrant
|
||||
|
||||
**Now it's time to try the code.** </br>
|
||||
Here’s a simplified Notebook to test ColPali for yourself:
|
||||
|
||||
[](https://colab.research.google.com/github/sabrinaaquino/colpali-qdrant-demo/blob/main/colpali_demo_binary.ipynb)
|
||||
|
||||
Our goal is to go through a dataset of multilingual newspaper articles like the ones below. We will detect which images contain text about **UFO's** and **Top Secret** events.
|
||||
|
||||

|
||||
|
||||
*The full dataset is accessible from the notebook.*
|
||||
|
||||
### Procedure
|
||||
|
||||
1. **Setup ColPali and Qdrant:** Import the necessary libraries, including a fine-tuned model optimized for your dataset (in this case, a UFO document set).
|
||||
2. **Dataset Preparation:** Load your document images into ColPali, previewing complex images to appreciate the challenge for traditional retrieval methods.
|
||||
3. **Qdrant Configuration:** Define your Qdrant collection, setting vector dimensions to 128. Enable Binary Quantization to optimize memory usage.
|
||||
4. **Batch Uploading Vectors:** Use a retry checkpoint to handle any exceptions during indexing. Batch processing allows you to adjust batch size based on available GPU resources.
|
||||
5. **Query Processing and Search:** Encode queries as multivectors for Qdrant. Set up rescoring and oversampling to fine-tune accuracy while optimizing speed.
|
||||
|
||||
### Results
|
||||
|
||||
> Success! Tests shows that search time is 2x faster than with Scalar Quantization.
|
||||
|
||||
This is significantly faster than with Scalar Quantization, and we still retrieved the top document matches with remarkable accuracy.
|
||||
|
||||
However, keep in mind that this is just a quick experiment. Performance may vary, so it's important to test Binary Quantization on your own datasets to see how it performs for your specific use case.
|
||||
|
||||
That said, it's promising to see Binary Quantization maintaining search quality while potentially offering performance improvements with ColPali.
|
||||
|
||||
## Future Directions with ColPali
|
||||

|
||||
|
||||
ColPali offers a promising, streamlined approach to document retrieval, especially for visually rich, complex documents. Its integration with Qdrant enables efficient large-scale vector storage and retrieval, ideal for machine learning applications requiring sophisticated document understanding.
|
||||
|
||||
If you’re interested in trying ColPali on your own datasets, join our [**vector search community on Discord**](https://qdrant.to/discord) for discussions, tutorials, and more insights into advanced document retrieval methods. Let us know in how you’re using ColPali or what applications you envision for it!
|
||||
|
||||
Thank you for reading, and stay tuned for more insights on vector search!
|
||||
|
||||
**References:**
|
||||
|
||||
[1] Faysse, M., Sibille, H., Wu, T., Omrani, B., Viaud, G., Hudelot, C., Colombo, P. (2024). **ColPali: Efficient Document Retrieval with Vision Language Models.** arXiv. https://doi.org/10.48550/arXiv.2407.01449
|
||||
|
||||
[2] van Strien, D. (2024). **Using ColPali with Qdrant to index and search a UFO document dataset.** Published October 2, 2024. Blog post: https://danielvanstrien.xyz/posts/post-with-code/colpali-qdrant/2024-10-02_using_colpali_with_qdrant.html
|
||||
|
||||
[3] Kacper Łukawski (2024). **Any Embedding Model Can Become a Late Interaction Model... If You Give It a Chance!** Qdrant Blog, August 14, 2024. Available at: https://qdrant.tech/articles/late-interaction-models/
|
||||
@@ -0,0 +1,232 @@
|
||||
---
|
||||
title: "Best Practices in RAG Evaluation: A Comprehensive Guide"
|
||||
draft: false
|
||||
short_description: "Learn how to evaluate RAG systems for accuracy and quality."
|
||||
description: "Explore best practices for evaluating Retrieval-Augmented Generation (RAG) systems. "
|
||||
preview_image: /blog/rag-evaluation-guide/social_preview.png
|
||||
social_preview_image: /blog/rag-evaluation-guide/social_preview.png
|
||||
date: 2024-11-24T00:00:00-08:00
|
||||
author: David Myriel
|
||||
featured: false
|
||||
tags:
|
||||
- RAG evaluation
|
||||
- retrieval-augmented generation
|
||||
- vector databases
|
||||
- LLM performance
|
||||
- semantic search
|
||||
- data retrieval
|
||||
- hybrid search
|
||||
- embeddings
|
||||
---
|
||||
|
||||
## Introduction
|
||||
|
||||
This guide will teach you how to evaluate a RAG system for both **accuracy** and **quality**. You will learn to maintain RAG performance by testing for search precision, recall, contextual relevance, and response accuracy.
|
||||
|
||||
**Building a RAG application is just the beginning;** it is crucial to test its usefulness for the end-user and calibrate its components for long-term stability.
|
||||
|
||||
RAG systems can encounter errors at any of the three crucial stages: retrieving relevant information, augmenting that information, and generating the final response. By systematically assessing and fine-tuning each component, you will be able to maintain a reliable and contextually relevant GenAI application that meets user needs.
|
||||
|
||||
## Why evaluate your RAG application?
|
||||
|
||||
### To avoid hallucinations and wrong answers
|
||||
|
||||

|
||||
|
||||
In the generation phase, hallucination is a notable issue where the LLM overlooks the context and fabricates information. This can lead to responses that are not grounded in reality.
|
||||
|
||||
Additionally, the generation of biased answers is a concern, as responses produced by the LLM can sometimes be harmful, inappropriate, or carry an inappropriate tone, thus posing risks in various applications and interactions.
|
||||
|
||||
### To enrich context augmented to your LLM
|
||||
|
||||
Augmentation processes face challenges such as outdated information, where responses may include data that is no longer current. Another issue is the presence of contextual gaps, where there is a lack of relational context between the retrieved documents.
|
||||
|
||||
> These gaps can result in incomplete or fragmented information being presented, reducing the overall coherence and relevance of the augmented responses.
|
||||
|
||||
### To maximize the search & retrieval process
|
||||
|
||||
When it comes to retrieval, one significant issue with search is the lack of precision, where not all documents retrieved are relevant to the query. This problem is compounded by poor recall, meaning not all relevant documents are successfully retrieved.
|
||||
|
||||
Additionally, the [“Lost in the Middle”](https://arxiv.org/abs/2307.03172) problem indicates that some LLMs may struggle with long contexts, particularly when crucial information is positioned in the middle of the document, leading to incomplete or less useful results.
|
||||
|
||||
## Recommended frameworks
|
||||
|
||||

|
||||
|
||||
To simplify the evaluation process, several powerful frameworks are available. Below we will explore three popular ones: **Ragas, Quotient AI, and Arize Phoenix**.
|
||||
|
||||
### Ragas: Testing RAG with questions and answers
|
||||
|
||||
[Ragas](https://docs.ragas.io/en/v0.0.17/index.html) (or RAG Assessment) uses a dataset of questions, ideal answers, and relevant context to compare a RAG system's generated answers with the ground truth. It provides metrics like faithfulness, relevance, and semantic similarity to assess retrieval and answer quality.
|
||||
|
||||
**Figure 1:** *Output of the Ragas framework, showcasing metrics like faithfulness, answer relevancy, context recall, precision, relevancy, entity recall, and answer similarity. These are used to evaluate the quality of RAG system responses.*
|
||||
|
||||

|
||||
|
||||
### Quotient: evaluating RAG pipelines with custom datasets
|
||||
|
||||
Quotient AI is another platform designed to streamline the evaluation of RAG systems. Developers can upload evaluation datasets as benchmarks to test different prompts and LLMs. These tests run as asynchronous jobs: Quotient AI automatically runs the RAG pipeline, generates responses and provides detailed metrics on faithfulness, relevance, and semantic similarity. The platform's full capabilities are accessible via a Python SDK, enabling you to access, analyze, and visualize your Quotient evaluation results to discover areas for improvement.
|
||||
|
||||
**Figure 2:** *Output of the Quotient framework, with statistics that define whether the dataset is properly manipulated throughout all stages of the RAG pipeline: indexing, chunking, search and context relevance.*
|
||||
|
||||

|
||||
|
||||
### Arize Phoenix: Visually Deconstructing Response Generation
|
||||
|
||||
[Arize Phoenix](https://docs.arize.com/phoenix) is an open-source tool that helps improve the performance of RAG systems by tracking how a response is built step-by-step. You can see these steps visually in Phoenix, which helps identify slowdowns and errors. You can define "[evaluators](https://docs.arize.com/phoenix/evaluation/concepts-evals/evaluation)" that use LLMs to assess the quality of outputs, detect hallucinations, and check answer accuracy. Phoenix also calculates key metrics like latency, token usage, and errors, giving you an idea of how efficiently your RAG system is working.
|
||||
|
||||
**Figure 3:** *The Arize Phoenix tool is intuitive to use and shows the entire process architecture as well as the steps that take place inside of retrieval, context and generation.*
|
||||
|
||||

|
||||
|
||||
## Why your RAG system might be underperforming
|
||||
|
||||

|
||||
|
||||
### You improperly ingested data to the vector database
|
||||
|
||||
Improper data ingestion can cause the loss of important contextual information, which is critical for generating accurate and coherent responses. Also, inconsistent data ingestion can cause the system to produce unreliable and inconsistent responses, undermining user trust and satisfaction.
|
||||
|
||||
Vector databases support different [indexing](https://qdrant.tech/documentation/concepts/indexing/) techniques. In order to know if you are ingesting data properly, you should always check how changes in variables related to indexing techniques affect data ingestion.
|
||||
|
||||
#### Solution: Pay attention to how your data is chunked
|
||||
|
||||
**Calibrate document chunk size:** The chunk size determines data granularity and impacts precision, recall, and relevance. It should be aligned with the token limit of the embedding model.
|
||||
|
||||
**Ensure proper chunk overlap:** This helps retain context by sharing data points across chunks. It should be managed with strategies like deduplication and content normalization.
|
||||
|
||||
**Develop a proper chunking/text splitting strategy**: Make sure your chunking/text splitting strategy is tailored to your on data type (e.g., HTML, markdown, code, PDF) and use-case nuances. For example, legal documents may be split by headings and subsections, and medical literature by sentence boundaries or key concepts.
|
||||
|
||||
**Figure 4:** *You can use utilities like [ChunkViz](https://chunkviz.up.railway.app/) to visualize different chunk splitting strategies, chunk sizes, and chunk overlaps.*
|
||||
|
||||

|
||||
|
||||
### You might be embedding data incorrectly
|
||||
|
||||
You want to ensure that the embedding model accurately understands and represents the data. If the generated embeddings are accurate, similar data points will be closely positioned in the vector space. The quality of an embedding model is typically measured using benchmarks like the [Massive Text Embedding Benchmark (MTEB)](https://huggingface.co/spaces/mteb/leaderboard), where the model’s output is compared against a ground-truth dataset.
|
||||
|
||||
#### Solution: Pick the right embedding model
|
||||
|
||||
The embedding model plays a critical role in capturing semantic relationships in data.
|
||||
|
||||
There are several embedding models you can choose from, and the [Massive Text Embedding Benchmark (MTEB) Leaderboard](https://huggingface.co/spaces/mteb/leaderboard) is a great resource for reference. Lightweight libraries like [FastEmbed](https://github.com/qdrant/fastembed) support the generation of vector embeddings using [popular text embedding models](https://qdrant.github.io/fastembed/examples/Supported_Models/#supported-text-embedding-models).
|
||||
|
||||
When choosing an embedding model, **retrieval performance** and **domain** **specificity**. You need to ensure that the model can capture semantic nuances, which affects the retrieval performance. For specialized domains, you may need to select or train a custom embedding model.
|
||||
|
||||
### Your retrieval procedure isn’t optimized
|
||||
|
||||

|
||||
|
||||
Semantic retrieval evaluation tests the effectiveness of your data retrieval. There are several metrics you can choose from:
|
||||
|
||||
- **Precision@k**: Measures the number of relevant documents in the top-k search results.
|
||||
- **Mean reciprocal rank (MRR)**: Considers the position of the first relevant document in the search results.
|
||||
- **Discounted cumulative gain (DCG) and normalized DCG (NDCG)**: Based on the relevance score of the documents.
|
||||
|
||||
By evaluating the retrieval quality using these metrics, you can assess the effectiveness of your retrieval step. For evaluating the ANN algorithm specifically, Precision@k is the most appropriate metric, as it directly measures how well the algorithm approximates exact search results.
|
||||
|
||||
#### Solution: Choose the best retrieval algorithm
|
||||
|
||||
Each new LLM with a larger context window claims to render RAG obsolete. However, studies like "[Lost in the Middle](https://arxiv.org/abs/2307.03172)" demonstrate that feeding entire documents to LLMs can diminish their ability to answer questions effectively. Therefore, the retrieval algorithm is crucial for fetching the most relevant data in the RAG system.
|
||||
|
||||
**Configure dense vector retrieval:** You need to choose the right [similarity metric](https://qdrant.tech/documentation/concepts/search/) to get the best retrieval quality. Metrics used in dense vector retrieval include Cosine Similarity, Dot Product, Euclidean Distance, and Manhattan Distance.
|
||||
|
||||
**Use sparse vectors & hybrid search where needed**: For sparse vectors, the algorithm choice of BM-25, SPLADE, or BM-42 will affect retrieval quality. Hybrid Search combines dense vector retrieval with sparse vector-based search.
|
||||
|
||||
**Leverage simple filtering:** This approach combines dense vector search with attribute filtering to narrow down the search results.
|
||||
|
||||
**Set correct hyperparameters:** Your Chunking Strategy, Chunk Size, Overlap, and Retrieval Window Size significantly impact the retrieval step and must be tailored to specific requirements.
|
||||
|
||||
**Introduce re-ranking:** Such methods may use cross-encoder models to re-score the results returned by vector search. Re-ranking can significantly improve retrieval and thus RAG system performance.
|
||||
|
||||
### LLM generation performance is suboptimal
|
||||
|
||||
The LLM is responsible for generating responses based on the retrieved context. The choice of LLM ranges from OpenAI’s GPT models to open-weight models. The LLM you choose will significantly influence the performance of a RAG system. Here are some areas to watch out for:
|
||||
|
||||
- **Response quality**: The LLM selection will influence the fluency, coherence, and factual accuracy of generated responses.
|
||||
- **System performance**: Inference speeds vary between LLMs. Slower inference speeds can impact response times.
|
||||
- **Domain knowledge**: For domain-specific RAG applications, you may need LLMs trained on that domain. Some LLMs are easier to fine-tune than others.
|
||||
|
||||
#### Solution: Test and Critically Analyze LLM Quality
|
||||
|
||||
The [Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard) can help guide your LLM selection. On this leaderboard, LLMs are ranked based on their scores on various benchmarks, such as IFEval, GPQA, MMLU-PRO, and others.
|
||||
|
||||
Evaluating LLMs involves several key metrics and methods. You can use these metrics or frameworks to evaluate if the LLM is delivering high-quality, relevant, and reliable responses.
|
||||
|
||||
**Table 1:** Methods for Measuring LLM Response Quality
|
||||
|
||||
| **Column header** | **Column header** |
|
||||
| --- | --- |
|
||||
| [Perplexity](https://huggingface.co/spaces/evaluate-metric/perplexity) | Measure how well the model predicts text. |
|
||||
| Human Evaluation | Rate responses based on relevance, coherence and quality. |
|
||||
| [BLEU](https://en.wikipedia.org/wiki/BLEU) | Used in translation tasks to compare generated output with reference translations. Higher scores (0-1) indicate better performance. |
|
||||
| [ROUGE](https://en.wikipedia.org/wiki/ROUGE_(metric)) | Evaluates summary quality by comparing generated summaries with reference summaries, calculating precision, recall and F1-score. |
|
||||
| [EleutherAI](https://github.com/EleutherAI/lm-evaluation-harness) | A framework to test LLMs on different evaluation tasks. |
|
||||
| [HELM](https://github.com/stanford-crfm/helm) | A framework to evaluate LLMs, focusing on 12 different aspects that are important in real-world model deployments. |
|
||||
| Diversity | Assesses the variety and uniqueness of responses, with higher scores indicating more diverse outputs. |
|
||||
|
||||
Many LLM evaluation frameworks offer flexibility to accommodate domain-specific or custom evaluations, addressing the key RAG metrics for your use case. These frameworks utilize either LLM-as-a-Judge or the [OpenAI Moderation API](https://platform.openai.com/docs/guides/moderation/overview) to ensure the moderation of responses from your AI applications.
|
||||
|
||||
## Working with custom datasets
|
||||
|
||||

|
||||
|
||||
First, create question and ground-truth answer pairs from source documents for the evaluation dataset. Ground-truth answers are the precise responses you expect from the RAG system. You can create these in multiple ways:
|
||||
|
||||
- **Hand-crafting your dataset:** Manually create questions and answers.
|
||||
- **Use LLM to create synthetic data:** Leverage LLMs like [T5](https://huggingface.co/docs/transformers/en/model_doc/t5) or OpenAI APIs.
|
||||
- **Use the Ragas framework**: [This method](https://docs.ragas.io/en/latest/getstarted/testset_generation.html) uses an LLM to generate various question types for evaluating RAG systems.
|
||||
- **Use FiddleCube**: [FiddleCube](https://www.fiddlecube.ai/) is a system that can help generate a range of question types aimed at different aspects of the testing process.
|
||||
|
||||
Once you have created a dataset, collect the retrieved context and the final answer generated by your RAG pipeline for each question.
|
||||
|
||||
**Figure 5:** *Here is an example of four evaluation metrics:*
|
||||
|
||||
- **question**: A set of questions based on the source document.
|
||||
- **ground_truth**: The anticipated accurate answers to the queries.
|
||||
- **context**: The context retrieved by the RAG pipeline for each query.
|
||||
- **answer**: The answer generated by the RAG pipeline for each query.
|
||||
|
||||

|
||||
|
||||
## Conclusion: What to look for when running tests
|
||||
|
||||
To understand if a RAG system is functioning as it should, you want to ensure:
|
||||
|
||||
- **Retrieval effectiveness**: The information retrieved is semantically relevant.
|
||||
- **Relevance of responses:** The generated response is meaningful.
|
||||
- **Coherence of generated responses**: The response is coherent and logically connected.
|
||||
- **Up-to-date responses**: The response is based on current data.
|
||||
|
||||
### Evaluating a RAG application from End-to-End (E2E)
|
||||
|
||||
The End-to-End (E2E) evaluation assesses the overall performance of the entire Retrieval-Augmented Generation (RAG) system. Here are some of the key factors you can measure:
|
||||
|
||||
- **Helpfulness**: Measures how well the system's responses assist users in achieving their goals.
|
||||
- **Groundedness**: Ensures that the responses are based on verifiable information from the retrieved context.
|
||||
- **Latency**: Monitors the response time of the system to ensure it meets the required speed and efficiency standards.
|
||||
- **Conciseness**: Evaluates whether the responses are brief yet comprehensive.
|
||||
- **Consistency**: Ensures that the system consistently delivers high-quality responses across different queries and contexts.
|
||||
|
||||
For instance, you can measure the quality of the generated responses with metrics like **Answer Semantic Similarity** and **Correctness**.
|
||||
|
||||
Measuring semantic similarity will tell you the difference between the generated answer and the ground truth, ranging from 0 to 1. This system cosine similarity to evaluate alignment in the vector space.
|
||||
|
||||
Checking answer correctness evaluates the overall agreement between the generated answer and the ground truth, combining factual correctness (measured by the F1 score) and answer similarity score.
|
||||
|
||||
### RAG evaluation is just the beginning
|
||||
|
||||

|
||||
|
||||
RAG evaluation is just the beginning. It lays the foundation for continuous improvement and long-term success of your system. Initially, it can help you identify and address immediate issues related to retrieval accuracy, contextual relevance, and response quality. However, as your RAG system evolves and is subjected to new data, use cases, and user interactions, you need to continue testing and calibrating.
|
||||
|
||||
By continuously evaluating your application, you can ensure that the system adapts to changing requirements and maintains its performance over time. You should regularly calibrate all components such as embedding models, retrieval algorithms, and the LLM itself. This iterative process will help you identify and fix emerging problems, optimize system parameters, and incorporate user feedback.
|
||||
|
||||
The practice of RAG evaluation is in the early stages of development. Keep this guide and wait as more techniques, models, and evaluation frameworks are developed. We strongly recommend you incorporate them into your evaluation process.
|
||||
|
||||
### Helpful Links
|
||||
- Join our community on [Discord](https://discord.com/invite/qdrant)
|
||||
- Check out our [latest articles](https://qdrant.tech/articles/)
|
||||
- Try a [free Qdrant cluster](https://cloud.qdrant.io/login)
|
||||
- Choose the right deployment option for your application. [Talk to sales](https://qdrant.tech/contact-us/)
|
||||
|
||||
@@ -2,10 +2,77 @@
|
||||
title: Home
|
||||
weight: 2
|
||||
hideTOC: true
|
||||
breadcrumb: false
|
||||
content:
|
||||
- partial: "documentation/banners/banner-a"
|
||||
title: Qdrant Documentation
|
||||
description: Qdrant is an AI-native vector database and a semantic search engine. You can use it to extract meaningful information from unstructured data.
|
||||
linkDescription: <a href="https://github.com/qdrant/qdrant_demo/" target="_blank">Clone this repo now</a> and build a search engine in five minutes.
|
||||
cloudButton:
|
||||
text: Cloud Quickstart
|
||||
url: /documentation/quickstart-cloud/
|
||||
localButton:
|
||||
text: Local Quickstart
|
||||
url: /documentation/quickstart/
|
||||
contained: true
|
||||
- partial: documentation/banners/banner-d
|
||||
developingTitle: Ready to start developing?
|
||||
developingDescription: Qdrant is open-source and can be self-hosted. However, the quickest way to get started is with our <a href="https://qdrant.to/cloud" target="_blank">free tier</a> on Qdrant Cloud. It scales easily and provides a UI where you can interact with data.
|
||||
developingBlock:
|
||||
title: Create your first Qdrant Cloud cluster today
|
||||
button:
|
||||
text: Get Started
|
||||
url: https://qdrant.to/cloud
|
||||
image:
|
||||
src: /img/rocket.svg
|
||||
alt: Rocket
|
||||
- partial: documentation/sections/cards-section
|
||||
title: Optimize Qdrant's performance
|
||||
description: Boost search speed, reduce latency, and improve the accuracy and memory usage of your Qdrant deployment.
|
||||
button:
|
||||
text: Learn More
|
||||
url: /documentation/guides/optimize/
|
||||
cardsPartial: documentation/cards/docs-cards
|
||||
cards:
|
||||
- id: 1
|
||||
tag: Documents
|
||||
icon:
|
||||
src: /icons/outline/documentation-blue.svg
|
||||
alt: Documents
|
||||
title: Distributed Deployment
|
||||
description: Scale Qdrant beyond a single node and optimize for high availability, fault tolerance, and billion-scale performance.
|
||||
link:
|
||||
url: /documentation/guides/distributed_deployment/
|
||||
text: Read More
|
||||
- id: 2
|
||||
tag: Documents
|
||||
icon:
|
||||
src: /icons/outline/documentation-blue.svg
|
||||
alt: Documents
|
||||
title: Multitenancy
|
||||
description: Build vector search apps that serve millions of users. Learn about data isolation, security, and performance tuning.
|
||||
link:
|
||||
url: /documentation/guides/multiple-partitions/
|
||||
text: Read More
|
||||
- id: 3
|
||||
tag: Blog
|
||||
tagColor: violet
|
||||
icon:
|
||||
src: /icons/outline/blog-purple.svg
|
||||
alt: Blog
|
||||
title: Vector Quantization
|
||||
description: Learn about cutting-edge techniques for vector quantization and how they can be used to improve search performance.
|
||||
link:
|
||||
url: /articles/what-is-vector-quantization/
|
||||
text: Read More
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
THIS CONTENT IS GOING TO BE IGNORED FOR NOW
|
||||
|
||||
# Documentation
|
||||
|
||||
Qdrant is an AI-native vector dabatase and a semantic search engine. You can use it to extract meaningful information from unstructured data. Want to see how it works? [Clone this repo now](https://github.com/qdrant/qdrant_demo/) and build a search engine in five minutes.
|
||||
Qdrant is an AI-native vector database and a semantic search engine. You can use it to extract meaningful information from unstructured data. Want to see how it works? [Clone this repo now](https://github.com/qdrant/qdrant_demo/) and build a search engine in five minutes.
|
||||
|
||||
|||
|
||||
|-:|:-|
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
---
|
||||
title: Advanced Retrieval
|
||||
weight: 17
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Advanced Tutorials
|
||||
|
||||
| |
|
||||
|----------------------------------------------------------|
|
||||
| [Use Collaborative Filtering to Build a Movie Recommendation System with Qdrant](/documentation/advanced-tutorials/collaborative-filtering/) |
|
||||
| [Build a Text/Image Multimodal Search System with Qdrant and FastEmbed](/documentation/advanced-tutorials/multimodal-search-fastembed/) |
|
||||
| [Navigate Your Codebase with Semantic Search and Qdrant](/documentation/advanced-tutorials/code-search/) |
|
||||
+5
-3
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Semantic code search
|
||||
weight: 22
|
||||
title: Search Through Your Codebase
|
||||
aliases:
|
||||
- /documentation/tutorials/code-search/
|
||||
weight: 2
|
||||
---
|
||||
|
||||
# Use semantic search to navigate your codebase
|
||||
# Navigate Your Codebase with Semantic Search and Qdrant
|
||||
|
||||
| Time: 45 min | Level: Intermediate | [](https://colab.research.google.com/github/qdrant/examples/blob/master/code-search/code-search.ipynb) | |
|
||||
|--------------|---------------------|--|----|
|
||||
+5
-3
@@ -1,13 +1,15 @@
|
||||
---
|
||||
title: Collaborative filtering
|
||||
title: Build a Recommendation System with Collaborative Filtering
|
||||
aliases:
|
||||
- /documentation/tutorials/collaborative-filtering/
|
||||
short_description: "Build an effective movie recommendation system using collaborative filtering and Qdrant's similarity search."
|
||||
description: "Build an effective movie recommendation system using collaborative filtering and Qdrant's similarity search."
|
||||
preview_image: /blog/collaborative-filtering/social_preview.png
|
||||
social_preview_image: /blog/collaborative-filtering/social_preview.png
|
||||
weight: 23
|
||||
weight: 3
|
||||
---
|
||||
|
||||
# Create a collaborative filtering system
|
||||
# Use Collaborative Filtering to Build a Movie Recommendation System with Qdrant
|
||||
|
||||
| Time: 45 min | Level: Intermediate | [](https://githubtocolab.com/qdrant/examples/blob/master/collaborative-filtering/collaborative-filtering.ipynb) | |
|
||||
|--------------|---------------------|--|----|
|
||||
+6
-4
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Multimodal Search
|
||||
weight: 4
|
||||
title: Setup Text/Image Multimodal Search
|
||||
aliases:
|
||||
- /documentation/tutorials/multimodal-search-fastembed/
|
||||
weight: 1
|
||||
---
|
||||
|
||||
# Multimodal Search with Qdrant and FastEmbed
|
||||
# Build a Multimodal Search System with Qdrant and FastEmbed
|
||||
|
||||
| Time: 15 min | Level: Beginner |Output: [GitHub](https://github.com/qdrant/examples/blob/master/multimodal-search/Multimodal_Search_with_FastEmbed.ipynb)|[](https://githubtocolab.com/qdrant/examples/blob/master/multimodal-search/Multimodal_Search_with_FastEmbed.ipynb) |
|
||||
| --- | ----------- | ----------- | ----------- |
|
||||
@@ -159,7 +161,7 @@ Image.open(client.search(
|
||||
|
||||

|
||||
|
||||
<h3 style="font-size: 1.25em;">Image-to-Text</h3>
|
||||
### Image-to-Text
|
||||
Now, let's do a reverse search with an image:
|
||||
|
||||
|
||||
@@ -0,0 +1,386 @@
|
||||
---
|
||||
title: Simple Agentic RAG System
|
||||
weight: 12
|
||||
partition: build
|
||||
social_preview_image: /documentation/examples/agentic-rag-crewai-zoom/social_preview.png
|
||||
---
|
||||

|
||||
|
||||
# Agentic RAG With CrewAI & Qdrant Vector Database
|
||||
|
||||
| Time: 45 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/examples/tree/master/agentic_rag_zoom_crewai) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
By combining the power of Qdrant for vector search and CrewAI for orchestrating modular agents, you can build systems that don't just answer questions but analyze, interpret, and act.
|
||||
|
||||
Traditional RAG systems focus on fetching data and generating responses, but they lack the ability to reason deeply or handle multi-step processes. Agentic RAG solves this by combining:
|
||||
|
||||
In this tutorial, we'll walk you through building an Agentic RAG system step by step. By the end, you'll have a working framework for storing data in a Qdrant Vector Database and extracting insights using CrewAI agents in conjunction with Vector Search over your data.
|
||||
|
||||
We already built this app for you. [Clone this repository](https://github.com/qdrant/examples/tree/master/agentic_rag_zoom_crewai) and follow along with the tutorial.
|
||||
|
||||
## What You'll Build
|
||||
In this hands-on tutorial, we'll create a system that:
|
||||
|
||||
1. Uses Qdrant to store and retrieve meeting transcripts as vector embeddings
|
||||
2. Leverages CrewAI agents to analyze and summarize meeting data
|
||||
3. Presents insights in a simple Streamlit interface for easy interaction
|
||||
|
||||
This project demonstrates how to build a Vector Search powered Agentic workflow to extract insights from meeting recordings. By combining Qdrant's vector search capabilities with CrewAI agents, users can search through and analyze their own meeting content.
|
||||
|
||||
The application first converts the meeting transcript into vector embeddings and stores them in a Qdrant vector database. It then uses CrewAI agents to query the vector database and extract insights from the meeting content. Finally, it uses Anthropic Claude to generate natural language responses to user queries based on the extracted insights from the vector database.
|
||||
|
||||
### How Does It Work?
|
||||
|
||||
When you interact with the system, here's what happens behind the scenes:
|
||||
|
||||
First the user submits a query to the system. In this example, we want to find out the avergae length of Marketing meetings. Since one of the data points from the meetings is the duration of the meeting, the agent can calculate the average duration of the meetings by averaging the duration of all meetings with the keyword "Marketing" in the topic or content.
|
||||
|
||||

|
||||
|
||||
Next, the agent used the `search_meetings` tool to search the Qdrant vector database for the most semantically similar meeting points. We asked about Marketing meetings, so the agent searched the database with the search meeting tool for all meetings with the keyword "Marketing" in the topic or content.
|
||||
|
||||

|
||||
|
||||
|
||||
Next, the agent used the `calculator` tool to find the average duration of the meetings.
|
||||
|
||||

|
||||
|
||||
Finally, the agent used the `Information Synthesizer` tool to synthesize the analysis and present it in a natural language format.
|
||||
|
||||

|
||||
|
||||
The user sees the final output in a chat-like interface.
|
||||
|
||||

|
||||
|
||||
The user can then continue to interact with the system by asking more questions.
|
||||
|
||||
|
||||
### Architecture
|
||||
|
||||
The system is built on three main components:
|
||||
- **Qdrant Vector Database**: Stores meeting transcripts and summaries as vector embeddings, enabling semantic search
|
||||
- **CrewAI Framework**: Coordinates AI agents that handle different aspects of meeting analysis
|
||||
- **Anthropic Claude**: Provides natural language understanding and response generation
|
||||
|
||||
1. **Data Processing Pipeline**
|
||||
- Processes meeting transcripts and metadata
|
||||
- Creates embeddings with SentenceTransformer
|
||||
- Manages Qdrant collection and data upload
|
||||
|
||||
2. **AI Agent System**
|
||||
- Implements CrewAI agent logic
|
||||
- Handles vector search integration
|
||||
- Processes queries with Claude
|
||||
|
||||
3. **User Interface**
|
||||
- Provides chat-like web interface
|
||||
- Shows real-time processing feedback
|
||||
- Maintains conversation history
|
||||
|
||||
---
|
||||
|
||||
## Getting Started
|
||||
|
||||

|
||||
|
||||
1. **Get API Credentials for Qdrant**:
|
||||
- Sign up for an account at [Qdrant Cloud](https://cloud.qdrant.io/).
|
||||
- Create a new cluster and copy the **Cluster URL** (format: https://xxx.gcp.cloud.qdrant.io).
|
||||
- Go to **Data Access Control** and generate an **API key**.
|
||||
|
||||
2. **Get API Credentials for AI Services**:
|
||||
- Get an API key from [Anthropic](https://www.anthropic.com/)
|
||||
- Get an API key from [OpenAI](https://platform.openai.com/)
|
||||
|
||||
---
|
||||
|
||||
## Setup
|
||||
|
||||
1. **Clone the Repository**:
|
||||
```bash
|
||||
git clone https://github.com/qdrant/examples.git
|
||||
cd agentic_rag_zoom_crewai
|
||||
```
|
||||
|
||||
2. **Create and Activate a Python Virtual Environment with Python 3.10 for compatibility**:
|
||||
```bash
|
||||
python3.10 -m venv venv
|
||||
source venv/bin/activate # Windows: venv\Scripts\activate
|
||||
```
|
||||
|
||||
3. **Install Dependencies**:
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
4. **Configure Environment Variables**:
|
||||
Create a `.env.local` file with:
|
||||
|
||||
```bash
|
||||
openai_api_key=your_openai_key_here
|
||||
anthropic_api_key=your_anthropic_key_here
|
||||
qdrant_url=your_qdrant_url_here
|
||||
qdrant_api_key=your_qdrant_api_key_here
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
### 1. Process Meeting Data
|
||||
The [`data_loader.py`](https://github.com/qdrant/examples/blob/master/agentic_rag_zoom_crewai/vector/data_loader.py) script processes meeting data and stores it in Qdrant:
|
||||
|
||||
```bash
|
||||
python vector/data_loader.py
|
||||
```
|
||||
|
||||
After this script has run, you should see a new collection in your Qdrant Cloud account called `zoom_recordings`. This collection contains the vector embeddings of the meeting transcripts. The points in the collection contain the original meeting data, including the topic, content, and summary.
|
||||
|
||||
### 2. Launch the Interface
|
||||
The [`streamlit_app.py`](https://github.com/qdrant/examples/blob/master/agentic_rag_zoom_crewai/vector/streamlit_app.py) is located in the `vector` folder. To launch it, run:
|
||||
|
||||
```bash
|
||||
streamlit run vector/streamlit_app.py
|
||||
```
|
||||
When you run this script, you will be able to interact with the system through a chat-like interface. Ask questions about the meeting content, and the system will use the AI agents to find the most relevant information and present it in a natural language format.
|
||||
|
||||
|
||||
### The Data Pipeline
|
||||
|
||||
At the heart of our system is the data processing pipeline:
|
||||
|
||||
```python
|
||||
class MeetingData:
|
||||
def _initialize(self):
|
||||
self.data_dir = Path(__file__).parent.parent / 'data'
|
||||
self.meetings = self._load_meetings()
|
||||
|
||||
self.qdrant_client = QdrantClient(
|
||||
url=os.getenv('qdrant_url'),
|
||||
api_key=os.getenv('qdrant_api_key')
|
||||
)
|
||||
self.embedding_model = SentenceTransformer('all-MiniLM-L6-v2')
|
||||
```
|
||||
The singleton pattern in data_loader.py is implemented through a MeetingData class that uses Python's __new__ and __init__ methods. The class maintains a private _instance variable to track if an instance exists, and a _initialized flag to ensure the initialization code only runs once. When creating a new instance with MeetingData(), __new__ first checks if _instance exists - if it doesn't, it creates one and sets the initialization flag to False. The __init__ method then checks this flag, and if it's False, runs the initialization code and sets the flag to True. This ensures that all subsequent calls to MeetingData() return the same instance with the same initialized resources.
|
||||
|
||||
When processing meetings, we need to consider both the content and context. Each meeting gets converted into a rich text representation before being transformed into a vector:
|
||||
|
||||
```python
|
||||
text_to_embed = f"""
|
||||
Topic: {meeting.get('topic', '')}
|
||||
Content: {meeting.get('vtt_content', '')}
|
||||
Summary: {json.dumps(meeting.get('summary', {}))}
|
||||
"""
|
||||
```
|
||||
|
||||
This structured format ensures our vector embeddings capture the full context of each meeting. But processing meetings one at a time would be inefficient. Instead, we batch process our data:
|
||||
|
||||
```python
|
||||
batch_size = 100
|
||||
for i in range(0, len(points), batch_size):
|
||||
batch = points[i:i + batch_size]
|
||||
self.qdrant_client.upsert(
|
||||
collection_name='zoom_recordings',
|
||||
points=batch
|
||||
)
|
||||
```
|
||||
|
||||
### Building the AI Agent System
|
||||
|
||||
Our AI system uses a tool-based approach. Let's start with the simplest tool - a calculator for meeting statistics:
|
||||
|
||||
```python
|
||||
class CalculatorTool(BaseTool):
|
||||
name: str = "calculator"
|
||||
description: str = "Perform basic mathematical calculations"
|
||||
|
||||
def _run(self, a: int, b: int) -> dict:
|
||||
return {
|
||||
"addition": a + b,
|
||||
"multiplication": a * b
|
||||
}
|
||||
```
|
||||
|
||||
But the real power comes from our vector search integration. This tool converts natural language queries into vector representations and searches our meeting database:
|
||||
|
||||
```python
|
||||
class SearchMeetingsTool(BaseTool):
|
||||
def _run(self, query: str) -> List[Dict]:
|
||||
response = openai_client.embeddings.create(
|
||||
model="text-embedding-ada-002",
|
||||
input=query
|
||||
)
|
||||
query_vector = response.data[0].embedding
|
||||
|
||||
return self.qdrant_client.search(
|
||||
collection_name='zoom_recordings',
|
||||
query_vector=query_vector,
|
||||
limit=10
|
||||
)
|
||||
```
|
||||
|
||||
The search results then feed into our analysis tool, which uses Claude to provide deeper insights:
|
||||
|
||||
```python
|
||||
class MeetingAnalysisTool(BaseTool):
|
||||
def _run(self, meeting_data: dict) -> Dict:
|
||||
meetings_text = self._format_meetings(meeting_data)
|
||||
|
||||
message = client.messages.create(
|
||||
model="claude-3-sonnet-20240229",
|
||||
messages=[{
|
||||
"role": "user",
|
||||
"content": f"Analyze these meetings:\n\n{meetings_text}"
|
||||
}]
|
||||
)
|
||||
```
|
||||
|
||||
### Orchestrating the Workflow
|
||||
|
||||
The magic happens when we bring these tools together under our agent framework. We create two specialized agents:
|
||||
|
||||
```python
|
||||
researcher = Agent(
|
||||
role='Research Assistant',
|
||||
goal='Find and analyze relevant information',
|
||||
tools=[calculator, searcher, analyzer]
|
||||
)
|
||||
|
||||
synthesizer = Agent(
|
||||
role='Information Synthesizer',
|
||||
goal='Create comprehensive and clear responses'
|
||||
)
|
||||
```
|
||||
|
||||
These agents work together in a coordinated workflow. The researcher gathers and analyzes information, while the synthesizer creates clear, actionable responses. This separation of concerns allows each agent to focus on its strengths.
|
||||
|
||||
### Building the User Interface
|
||||
|
||||
The Streamlit interface provides a clean, chat-like experience for interacting with our AI system. Let's start with the basic setup:
|
||||
|
||||
```python
|
||||
st.set_page_config(
|
||||
page_title="Meeting Assistant",
|
||||
page_icon="🤖",
|
||||
layout="wide"
|
||||
)
|
||||
```
|
||||
|
||||
To make the interface more engaging, we add custom styling that makes the output easier to read:
|
||||
|
||||
```python
|
||||
st.markdown("""
|
||||
<style>
|
||||
.stApp {
|
||||
max-width: 1200px;
|
||||
margin: 0 auto;
|
||||
}
|
||||
.output-container {
|
||||
background-color: #f0f2f6;
|
||||
padding: 20px;
|
||||
border-radius: 10px;
|
||||
margin: 10px 0;
|
||||
}
|
||||
</style>
|
||||
""", unsafe_allow_html=True)
|
||||
```
|
||||
|
||||
One of the key features is real-time feedback during processing. We achieve this with a custom output handler:
|
||||
|
||||
```python
|
||||
class ConsoleOutput:
|
||||
def __init__(self, placeholder):
|
||||
self.placeholder = placeholder
|
||||
self.buffer = []
|
||||
self.update_interval = 0.5 # seconds
|
||||
self.last_update = time.time()
|
||||
|
||||
def write(self, text):
|
||||
self.buffer.append(text)
|
||||
if time.time() - self.last_update > self.update_interval:
|
||||
self._update_display()
|
||||
```
|
||||
|
||||
This handler buffers the output and updates the display periodically, creating a smooth user experience. When a user sends a query, we process it with visual feedback:
|
||||
|
||||
```python
|
||||
with st.chat_message("assistant"):
|
||||
message_placeholder = st.empty()
|
||||
progress_bar = st.progress(0)
|
||||
console_placeholder = st.empty()
|
||||
|
||||
try:
|
||||
console_output = ConsoleOutput(console_placeholder)
|
||||
with contextlib.redirect_stdout(console_output):
|
||||
progress_bar.progress(0.3)
|
||||
full_response = get_crew_response(prompt)
|
||||
progress_bar.progress(1.0)
|
||||
```
|
||||
|
||||
The interface maintains a chat history, making it feel like a natural conversation:
|
||||
|
||||
```python
|
||||
if "messages" not in st.session_state:
|
||||
st.session_state.messages = []
|
||||
|
||||
for message in st.session_state.messages:
|
||||
with st.chat_message(message["role"]):
|
||||
st.markdown(message["content"])
|
||||
```
|
||||
|
||||
We also include helpful examples and settings in the sidebar:
|
||||
|
||||
```python
|
||||
with st.sidebar:
|
||||
st.header("Settings")
|
||||
search_limit = st.slider("Number of results", 1, 10, 5)
|
||||
|
||||
analysis_depth = st.select_slider(
|
||||
"Analysis Depth",
|
||||
options=["Basic", "Standard", "Detailed"],
|
||||
value="Standard"
|
||||
)
|
||||
```
|
||||
|
||||
This combination of features creates an interface that's both powerful and approachable. Users can see their query being processed in real-time, adjust settings to their needs, and maintain context through the chat history.
|
||||
|
||||
---
|
||||
## Conclusion
|
||||
|
||||

|
||||
|
||||
This tutorial has demonstrated how to build a sophisticated meeting analysis system that combines vector search with AI agents. Let's recap the key components we've covered:
|
||||
|
||||
1. **Vector Search Integration**
|
||||
- Efficient storage and retrieval of meeting content using Qdrant
|
||||
- Semantic search capabilities through vector embeddings
|
||||
- Batched processing for optimal performance
|
||||
|
||||
2. **AI Agent Framework**
|
||||
- Tool-based approach for modular functionality
|
||||
- Specialized agents for research and analysis
|
||||
- Integration with Claude for intelligent insights
|
||||
|
||||
3. **Interactive Interface**
|
||||
- Real-time feedback and progress tracking
|
||||
- Persistent chat history
|
||||
- Configurable search and analysis settings
|
||||
|
||||
The resulting system demonstrates the power of combining vector search with AI agents to create an intelligent meeting assistant. By following this tutorial, you've learned how to:
|
||||
- Process and store meeting data efficiently
|
||||
- Implement semantic search capabilities
|
||||
- Create specialized AI agents for analysis
|
||||
- Build an intuitive user interface
|
||||
|
||||
This foundation can be extended in many ways, such as:
|
||||
- Adding more specialized agents
|
||||
- Implementing additional analysis tools
|
||||
- Enhancing the user interface
|
||||
- Integrating with other data sources
|
||||
|
||||
The code is available in the [repository](https://github.com/qdrant/examples/tree/master/agentic_rag_zoom_crewai), and we encourage you to experiment with your own modifications and improvements.
|
||||
|
||||
---
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
title: Vector Search Basics
|
||||
aliases:
|
||||
- /documentation/tutorials/
|
||||
weight: 16
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Beginner Tutorials
|
||||
|
||||
| |
|
||||
|----------------------------------------------------|
|
||||
| [Build Your First Semantic Search Engine in 5 Minutes](/documentation/beginner-tutorials/search-beginners/) |
|
||||
| [Build a Neural Search Service with Sentence Transformers and Qdrant](/documentation/beginner-tutorials/neural-search/) |
|
||||
| [Build a Hybrid Search Service with FastEmbed and Qdrant](/documentation/beginner-tutorials/hybrid-search-fastembed/) |
|
||||
| [Measure and Improve Retrieval Quality in Semantic Search](/documentation/beginner-tutorials/retrieval-quality/) |
|
||||
+4
-5
@@ -1,12 +1,11 @@
|
||||
---
|
||||
title: Hybrid Search with Fastembed
|
||||
weight: 2
|
||||
|
||||
title: Setup Hybrid Search with FastEmbed
|
||||
aliases:
|
||||
- /documentation/tutorials/neural-search-fastembed/
|
||||
- /documentation/tutorials/hybrid-search-fastembed/
|
||||
weight: 3
|
||||
---
|
||||
|
||||
# Create a Hybrid Search Service with Fastembed
|
||||
# Build a Hybrid Search Service with FastEmbed and Qdrant
|
||||
|
||||
| Time: 20 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/qdrant_demo/) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
+6
-4
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Neural Search Service
|
||||
weight: 1
|
||||
title: Build a Neural Search Service
|
||||
aliases:
|
||||
- /documentation/tutorials/neural-search/
|
||||
weight: 2
|
||||
---
|
||||
|
||||
# Create a Simple Neural Search Service
|
||||
# Build a Neural Search Service with Sentence Transformers and Qdrant
|
||||
|
||||
| Time: 30 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/qdrant_demo/tree/sentense-transformers) | [](https://colab.research.google.com/drive/1kPktoudAP8Tu8n8l-iVMOQhVmHkWV_L9?usp=sharing) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
@@ -14,7 +16,7 @@ A neural search service uses artificial neural networks to improve the accuracy
|
||||
|
||||
<aside role="status">
|
||||
There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers.
|
||||
Check it out <a href="/documentation/tutorials/hybrid-search-fastembed/">here</a>.
|
||||
Check it out <a href="/documentation/beginner-tutorials/hybrid-search-fastembed/">here</a>.
|
||||
</aside>
|
||||
|
||||
|
||||
+5
-3
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Measure retrieval quality
|
||||
weight: 21
|
||||
title: Measure Search Quality
|
||||
aliases:
|
||||
- /documentation/tutorials/retrieval-quality/
|
||||
weight: 4
|
||||
---
|
||||
|
||||
# Measure retrieval quality
|
||||
# Measure and Improve Retrieval Quality in Semantic Search
|
||||
|
||||
| Time: 30 min | Level: Intermediate | | |
|
||||
|--------------|---------------------|--|----|
|
||||
+3
-2
@@ -1,11 +1,12 @@
|
||||
---
|
||||
title: Semantic Search 101
|
||||
weight: -100
|
||||
weight: 1
|
||||
aliases:
|
||||
- /documentation/tutorials/mighty.md/
|
||||
- /documentation/tutorials/search-beginners/
|
||||
---
|
||||
|
||||
# Semantic Search for Beginners
|
||||
# Build Your First Semantic Search Engine in 5 Minutes
|
||||
|
||||
| Time: 5 - 15 min | Level: Beginner | | |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
@@ -0,0 +1,49 @@
|
||||
---
|
||||
title: Build World-Class Applications
|
||||
slug: build
|
||||
breadcrumb: false
|
||||
content:
|
||||
- partial: documentation/banners/banner-c
|
||||
title: Build World-Class Applications
|
||||
description: Dev-portal Build
|
||||
image:
|
||||
src: /img/dev-portal-build/spanish-ai-app-hero.png
|
||||
alt: Spanish AI app.png
|
||||
startedButton:
|
||||
text: Get Started
|
||||
url: https://qdrant.to/cloud
|
||||
- partial: documentation/sections/cards-section
|
||||
title: Start Building
|
||||
description: Deploy and manage high-performance vector search clusters across cloud environments. Easily scale with fully managed cloud solutions, integrate seamlessly across hybrid setups, or maintain complete control with private cloud deployments in Kubernetes.
|
||||
cardsPartial: documentation/cards/docs-cards
|
||||
cards:
|
||||
- id: 1
|
||||
image:
|
||||
src: /img/dev-portal-build/search.png
|
||||
alt: Search
|
||||
title: Search
|
||||
description: Build a simple neural search service with Qdrant and FastEmbed. Learn how to upload data, create indexes, and run search queries.
|
||||
link:
|
||||
url: /documentation/beginner-tutorials/hybrid-search-fastembed/
|
||||
text: Read More
|
||||
- id: 2
|
||||
image:
|
||||
src: /img/dev-portal-build/rag.png
|
||||
alt: RAG
|
||||
title: RAG
|
||||
description: Build end-to-end prototype chatbots. Learn how Qdrant integrates with popular RAG frameworks like LangChain and Llamaindex.
|
||||
link:
|
||||
url: /documentation/frameworks/langchain/
|
||||
text: Read More
|
||||
- id: 3
|
||||
image:
|
||||
src: /img/dev-portal-build/pipelines.png
|
||||
alt: Pipelines
|
||||
title: Pipelines
|
||||
description: Integrate Qdrant into your data infrastructure by connecting with popular data engineering tools.
|
||||
link:
|
||||
url: /documentation/send-data/
|
||||
text: Read More
|
||||
partition: build
|
||||
hideInSidebar: true
|
||||
---
|
||||
@@ -0,0 +1,73 @@
|
||||
---
|
||||
title: Welcome to Qdrant Cloud
|
||||
slug: cloud-intro
|
||||
breadcrumb: false
|
||||
content:
|
||||
- partial: documentation/banners/banner-b
|
||||
title: Welcome to Qdrant Cloud
|
||||
description: Dev-portal Cloud
|
||||
image:
|
||||
src: /img/dev-portal-cloud/dev-portal-cloud-hero.png
|
||||
alt: Qdrant cloud dashboard
|
||||
startedButton:
|
||||
text: Get Started
|
||||
url: https://qdrant.to/cloud
|
||||
- partial: documentation/sections/cards-section
|
||||
title: Managed Services
|
||||
description: Deploy and manage high-performance vector search clusters across cloud environments. Easily scale with fully managed cloud solutions, integrate seamlessly across hybrid setups, or maintain complete control with private cloud deployments in Kubernetes.
|
||||
cardsPartial: documentation/cards/docs-cards
|
||||
cards:
|
||||
- id: 1
|
||||
image:
|
||||
src: /img/dev-portal-cloud/managed-cloud.png
|
||||
alt: Qdrant Cloud
|
||||
title: Qdrant Cloud
|
||||
description: Qdrant Managed Cloud is our SaaS solution, providing managed Qdrant database clusters on the cloud.
|
||||
link:
|
||||
url: /documentation/cloud/
|
||||
text: Read More
|
||||
- id: 2
|
||||
image:
|
||||
src: /img/dev-portal-cloud/hybrid-cloud.png
|
||||
alt: Hybrid Cloud
|
||||
title: Hybrid Cloud
|
||||
description: Deploy and manage your vector database across diverse environments, ensuring performance, security, and cost efficiency.
|
||||
link:
|
||||
url: /documentation/hybrid-cloud/
|
||||
text: Read More
|
||||
- id: 3
|
||||
image:
|
||||
src: /img/dev-portal-cloud/private-cloud.png
|
||||
alt: Private Cloud
|
||||
title: Private Cloud
|
||||
description: Qdrant Private Cloud allows you to manage Qdrant database clusters in any Kubernetes cluster on any infrastructure.
|
||||
link:
|
||||
url: /documentation/private-cloud/
|
||||
text: Read More
|
||||
- partial: documentation/sections/cards-section
|
||||
title: Customer Support
|
||||
description: Stream, index, and migrate data to Qdrant with these essential tools and strategies.
|
||||
cardsPartial: documentation/cards/docs-cards
|
||||
cardsPerRow: 2
|
||||
cards:
|
||||
- id: 1
|
||||
icon:
|
||||
src: /icons/outline/discord-purple.svg
|
||||
alt: Discord icon
|
||||
title: Community Support
|
||||
description: Join 6,000+ active members to learn, collaborate, and participate in Qdrant’s latest activities.
|
||||
link:
|
||||
text: Join our Discord
|
||||
url: https://qdrant.to/discord
|
||||
- id: 2
|
||||
icon:
|
||||
src: /icons/outline/support-blue.svg
|
||||
alt: Support icon
|
||||
title: Qdrant Cloud Support
|
||||
description: Paying customers have access to our Support team. Links to the support portal are available in the Qdrant Cloud Console.
|
||||
link:
|
||||
text: Join Qdrant
|
||||
url: https://qdrant.to/cloud
|
||||
partition: cloud
|
||||
hideInSidebar: true
|
||||
---
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
title: Infrastructure Tools
|
||||
weight: 28
|
||||
partition: cloud
|
||||
---
|
||||
|
||||
## Cloud Tools
|
||||
|
||||
| Integration | Description |
|
||||
| ----------------------------------- | ------------------------------------------------------------------------------------------- |
|
||||
| [Pulumi](/documentation/cloud-tools/pulumi/) | Infrastructure as code tool for creating, deploying, and managing cloud infrastructure |
|
||||
| [Terraform](/documentation/cloud-tools/terraform/) | infrastructure as code tool to define resources in human-readable configuration files. |
|
||||
|
||||
+2
-1
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: Pulumi
|
||||
aliases: [ ../platforms/pulumi ]
|
||||
aliases:
|
||||
- /documentation/infrastructure/pulumi/
|
||||
---
|
||||
|
||||

|
||||
+2
-1
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: Terraform
|
||||
aliases: [ ../platforms/terraform ]
|
||||
aliases:
|
||||
- /documentation/infrastructure/terraform/
|
||||
---
|
||||
|
||||

|
||||
@@ -3,6 +3,7 @@ title: Managed Cloud
|
||||
weight: 12
|
||||
aliases:
|
||||
- /documentation/overview/qdrant-alternatives/documentation/cloud/
|
||||
partition: cloud
|
||||
---
|
||||
|
||||
# About Qdrant Managed Cloud
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
title: Concepts
|
||||
weight: 8
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Concepts
|
||||
|
||||
@@ -200,7 +200,7 @@ curl -X PUT http://localhost:6333/collections/{collection_name} \
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
@@ -633,7 +633,7 @@ And additionally, sparse vectors and dense vectors must have different names wit
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"sparse_vectors": {
|
||||
"text": { },
|
||||
"text": { }
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -656,6 +656,7 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config={},
|
||||
sparse_vectors_config={
|
||||
"text": models.SparseVectorParams(),
|
||||
},
|
||||
@@ -858,7 +859,7 @@ curl -X PATCH http://localhost:6333/collections/{collection_name} \
|
||||
```python
|
||||
client.update_collection(
|
||||
collection_name="{collection_name}",
|
||||
optimizer_config=models.OptimizersConfigDiff(indexing_threshold=10000),
|
||||
optimizers_config=models.OptimizersConfigDiff(indexing_threshold=10000),
|
||||
)
|
||||
```
|
||||
|
||||
@@ -931,7 +932,7 @@ The following parameters can be updated:
|
||||
* `optimizers_config` - see [optimizer](/documentation/concepts/optimizer/) for details.
|
||||
* `hnsw_config` - see [indexing](/documentation/concepts/indexing/#vector-index) for details.
|
||||
* `quantization_config` - see [quantization](/documentation/guides/quantization/#setting-up-quantization-in-qdrant) for details.
|
||||
* `vectors` - vector-specific configuration, including individual `hnsw_config`, `quantization_config` and `on_disk` settings.
|
||||
* `vectors_config` - vector-specific configuration, including individual `hnsw_config`, `quantization_config` and `on_disk` settings.
|
||||
* `params` - other collection parameters, including `write_consistency_factor` and `on_disk_payload`.
|
||||
|
||||
Full API specification is available in [schema definitions](https://api.qdrant.tech/api-reference/collections/update-collection).
|
||||
|
||||
@@ -980,6 +980,10 @@ discover_queries = [
|
||||
limit=10,
|
||||
),
|
||||
]
|
||||
|
||||
client.query_batch_points(
|
||||
collection_name="{collection_name}", requests=discover_queries
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
@@ -1189,6 +1193,10 @@ discover_queries = [
|
||||
limit=10,
|
||||
),
|
||||
]
|
||||
|
||||
client.query_batch_points(
|
||||
collection_name="{collection_name}", requests=discover_queries
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
@@ -1367,6 +1375,8 @@ POST /collections/{collection_name}/points/search/matrix/pairs
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search_matrix_pairs(
|
||||
collection_name="{collection_name}",
|
||||
sample=10,
|
||||
@@ -1395,7 +1405,7 @@ QdrantClient client =
|
||||
client
|
||||
.searchMatrixPairsAsync(
|
||||
Points.SearchMatrixPoints.newBuilder()
|
||||
.setCollectionName(collectionName)
|
||||
.setCollectionName("{collection_name}")
|
||||
.setFilter(Filter.newBuilder().addMust(matchKeyword("color", "red")).build())
|
||||
.setSample(10)
|
||||
.setLimit(2)
|
||||
@@ -1448,7 +1458,7 @@ using static Qdrant.Client.Grpc.Conditions;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.SearchMatrixPairs(
|
||||
await client.SearchMatrixPairsAsync(
|
||||
collectionName: "{collection_name}",
|
||||
filter: MatchKeyword("color", "red"),
|
||||
sample: 10,
|
||||
@@ -1533,6 +1543,8 @@ POST /collections/{collection_name}/points/search/matrix/offsets
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.search_matrix_offsets(
|
||||
collection_name="{collection_name}",
|
||||
sample=10,
|
||||
@@ -1561,7 +1573,7 @@ QdrantClient client =
|
||||
client
|
||||
.searchMatrixOffsetsAsync(
|
||||
SearchMatrixPoints.newBuilder()
|
||||
.setCollectionName(collectionName)
|
||||
.setCollectionName("{collection_name}")
|
||||
.setFilter(Filter.newBuilder().addMust(matchKeyword("color", "red")).build())
|
||||
.setSample(10)
|
||||
.setLimit(2)
|
||||
@@ -1614,7 +1626,7 @@ using static Qdrant.Client.Grpc.Conditions;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.SearchMatrixOffsets(
|
||||
await client.SearchMatrixOffsetsAsync(
|
||||
collectionName: "{collection_name}",
|
||||
filter: MatchKeyword("color", "red"),
|
||||
sample: 10,
|
||||
|
||||
@@ -449,7 +449,7 @@ Filtered points would be:
|
||||
]
|
||||
```
|
||||
|
||||
When using `must_not`, the clause becomes `true` if none if the conditions listed inside `should` is satisfied.
|
||||
When using `must_not`, the clause becomes `true` if none of the conditions listed inside `should` is satisfied.
|
||||
In this sense, `must_not` is equivalent to the expression `(NOT A) AND (NOT B) AND (NOT C)`.
|
||||
|
||||
### Clauses combination
|
||||
@@ -2158,7 +2158,7 @@ Functionally, it will work with `keyword` and `uuid` indexes exactly the same, b
|
||||
{
|
||||
"key": "uuid",
|
||||
"match": {
|
||||
"uuid": "f47ac10b-58cc-4372-a567-0e02b2c3d479"
|
||||
"value": "f47ac10b-58cc-4372-a567-0e02b2c3d479"
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -2166,14 +2166,14 @@ Functionally, it will work with `keyword` and `uuid` indexes exactly the same, b
|
||||
```python
|
||||
models.FieldCondition(
|
||||
key="uuid",
|
||||
match=models.MatchValue(uuid="f47ac10b-58cc-4372-a567-0e02b2c3d479"),
|
||||
match=models.MatchValue(value="f47ac10b-58cc-4372-a567-0e02b2c3d479"),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
{
|
||||
key: 'uuid',
|
||||
match: {uuid: 'f47ac10b-58cc-4372-a567-0e02b2c3d479'}
|
||||
match: {value: 'f47ac10b-58cc-4372-a567-0e02b2c3d479'}
|
||||
}
|
||||
```
|
||||
|
||||
@@ -2462,57 +2462,59 @@ models.FieldCondition(
|
||||
|
||||
```typescript
|
||||
{
|
||||
key: 'location',
|
||||
geo_polygon: {
|
||||
exterior: {
|
||||
points: [
|
||||
{
|
||||
lon: -70.0,
|
||||
lat: -70.0
|
||||
},
|
||||
{
|
||||
lon: 60.0,
|
||||
lat: -70.0
|
||||
},
|
||||
{
|
||||
lon: 60.0,
|
||||
lat: 60.0
|
||||
},
|
||||
{
|
||||
lon: -70.0,
|
||||
lat: 60.0
|
||||
},
|
||||
{
|
||||
lon: -70.0,
|
||||
lat: -70.0
|
||||
}
|
||||
]
|
||||
key: "location",
|
||||
geo_polygon: {
|
||||
exterior: {
|
||||
points: [
|
||||
{
|
||||
lon: -70.0,
|
||||
lat: -70.0
|
||||
},
|
||||
interiors: {
|
||||
points: [
|
||||
{
|
||||
lon: -65.0,
|
||||
lat: -65.0
|
||||
},
|
||||
{
|
||||
lon: 0.0,
|
||||
lat: -65.0
|
||||
},
|
||||
{
|
||||
lon: 0.0,
|
||||
lat: 0.0
|
||||
},
|
||||
{
|
||||
lon: -65.0,
|
||||
lat: 0.0
|
||||
},
|
||||
{
|
||||
lon: -65.0,
|
||||
lat: -65.0
|
||||
}
|
||||
]
|
||||
{
|
||||
lon: 60.0,
|
||||
lat: -70.0
|
||||
},
|
||||
{
|
||||
lon: 60.0,
|
||||
lat: 60.0
|
||||
},
|
||||
{
|
||||
lon: -70.0,
|
||||
lat: 60.0
|
||||
},
|
||||
{
|
||||
lon: -70.0,
|
||||
lat: -70.0
|
||||
}
|
||||
}
|
||||
]
|
||||
},
|
||||
interiors: [
|
||||
{
|
||||
points: [
|
||||
{
|
||||
lon: -65.0,
|
||||
lat: -65.0
|
||||
},
|
||||
{
|
||||
lon: 0,
|
||||
lat: -65.0
|
||||
},
|
||||
{
|
||||
lon: 0,
|
||||
lat: 0
|
||||
},
|
||||
{
|
||||
lon: -65.0,
|
||||
lat: 0
|
||||
},
|
||||
{
|
||||
lon: -65.0,
|
||||
lat: -65.0
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
@@ -2764,7 +2766,7 @@ models.IsEmptyCondition(
|
||||
```typescript
|
||||
{
|
||||
is_empty: {
|
||||
key: "reports";
|
||||
key: "reports"
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -2820,7 +2822,7 @@ models.IsNullCondition(
|
||||
```typescript
|
||||
{
|
||||
is_null: {
|
||||
key: "reports";
|
||||
key: "reports"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
@@ -89,7 +89,7 @@ client.query_points(
|
||||
limit=20,
|
||||
),
|
||||
models.Prefetch(
|
||||
query=[0.01, 0.45, 0.67, ...], # <-- dense vector
|
||||
query=[0.01, 0.45, 0.67], # <-- dense vector
|
||||
using="dense",
|
||||
limit=20,
|
||||
),
|
||||
@@ -284,7 +284,7 @@ client.query_points(
|
||||
using="mrl_byte",
|
||||
limit=1000,
|
||||
),
|
||||
query=[0.01, 0.299, 0.45, 0.67, ...], # <-- full vector
|
||||
query=[0.01, 0.299, 0.45, 0.67], # <-- full vector
|
||||
using="full",
|
||||
limit=10,
|
||||
)
|
||||
@@ -296,14 +296,14 @@ import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
prefetch: {
|
||||
query: [1, 23, 45, 67], // <------------- small byte vector
|
||||
using: 'mrl_byte',
|
||||
limit: 1000,
|
||||
},
|
||||
query: [0.01, 0.299, 0.45, 0.67, ...], // <-- full vector,
|
||||
using: 'full',
|
||||
limit: 10,
|
||||
prefetch: {
|
||||
query: [1, 23, 45, 67], // <------------- small byte vector
|
||||
using: 'mrl_byte',
|
||||
limit: 1000,
|
||||
},
|
||||
query: [0.01, 0.299, 0.45, 0.67], // <-- full vector,
|
||||
using: 'full',
|
||||
limit: 10,
|
||||
});
|
||||
```
|
||||
|
||||
@@ -428,13 +428,13 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
prefetch=models.Prefetch(
|
||||
query=[0.01, 0.45, 0.67, ...], # <-- dense vector
|
||||
query=[0.01, 0.45, 0.67, 0.53], # <-- dense vector
|
||||
limit=100,
|
||||
),
|
||||
query=[
|
||||
[0.1, 0.2, ...], # <─┐
|
||||
[0.2, 0.1, ...], # < ├─ multi-vector
|
||||
[0.8, 0.9, ...], # < ┘
|
||||
[0.1, 0.2, 0.32], # <─┐
|
||||
[0.2, 0.1, 0.52], # < ├─ multi-vector
|
||||
[0.8, 0.9, 0.93], # < ┘
|
||||
],
|
||||
using="colbert",
|
||||
limit=10,
|
||||
@@ -608,14 +608,14 @@ client.query_points(
|
||||
using="mrl_byte",
|
||||
limit=1000,
|
||||
),
|
||||
query=[0.01, 0.45, 0.67, ...], # <-- full dense vector
|
||||
query=[0.01, 0.45, 0.67], # <-- full dense vector
|
||||
using="full",
|
||||
limit=100,
|
||||
),
|
||||
query=[
|
||||
[0.1, 0.2, ...], # <─┐
|
||||
[0.2, 0.1, ...], # < ├─ multi-vector
|
||||
[0.8, 0.9, ...], # < ┘
|
||||
[0.17, 0.23, 0.52], # <─┐
|
||||
[0.22, 0.11, 0.63], # < ├─ multi-vector
|
||||
[0.86, 0.93, 0.12], # < ┘
|
||||
],
|
||||
using="colbert",
|
||||
limit=10,
|
||||
@@ -628,23 +628,23 @@ import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
prefetch: {
|
||||
prefetch: {
|
||||
prefetch: {
|
||||
query: [1, 23, 45, 67, ...], // <------------- small byte vector
|
||||
using: 'mrl_byte',
|
||||
limit: 1000,
|
||||
},
|
||||
query: [0.01, 0.45, 0.67, ...], // <-- full dense vector
|
||||
using: 'full',
|
||||
limit: 100,
|
||||
query: [1, 23, 45, 67], // <------------- small byte vector
|
||||
using: 'mrl_byte',
|
||||
limit: 1000,
|
||||
},
|
||||
query: [
|
||||
[0.1, 0.2], // <─┐
|
||||
[0.2, 0.1], // < ├─ multi-vector
|
||||
[0.8, 0.9], // < ┘
|
||||
],
|
||||
using: 'colbert',
|
||||
limit: 10,
|
||||
query: [0.01, 0.45, 0.67], // <-- full dense vector
|
||||
using: 'full',
|
||||
limit: 100,
|
||||
},
|
||||
query: [
|
||||
[0.1, 0.2], // <─┐
|
||||
[0.2, 0.1], // < ├─ multi-vector
|
||||
[0.8, 0.9], // < ┘
|
||||
],
|
||||
using: 'colbert',
|
||||
limit: 10,
|
||||
});
|
||||
```
|
||||
|
||||
@@ -832,7 +832,7 @@ let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(PointId::new("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")))
|
||||
.query(Query::new_nearest("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
@@ -913,7 +913,7 @@ client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query="43cf51e2-8777-4f52-bc74-c2cbde0c8b04", # <--- point id
|
||||
using="512d-vector",
|
||||
lookup_from=models.LookupFrom(
|
||||
lookup_from=models.LookupLocation(
|
||||
collection="another_collection", # <--- other collection name
|
||||
vector="image-512", # <--- vector name in the other collection
|
||||
)
|
||||
@@ -943,7 +943,7 @@ let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(PointId::new("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")))
|
||||
.query(Query::new_nearest("43cf51e2-8777-4f52-bc74-c2cbde0c8b04"))
|
||||
.using("512d-vector")
|
||||
.lookup_from(
|
||||
LookupLocationBuilder::new("another_collection")
|
||||
@@ -1079,21 +1079,21 @@ client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
prefetch=[
|
||||
models.Prefetch(
|
||||
query=[0.01, 0.45, 0.67, ...], # <-- dense vector
|
||||
query=[0.01, 0.45, 0.67], # <-- dense vector
|
||||
filter=models.Filter(
|
||||
must=models.FieldCondition(
|
||||
key="color",
|
||||
match=models.Match(value="red"),
|
||||
match=models.MatchValue(value="red"),
|
||||
),
|
||||
),
|
||||
limit=10,
|
||||
),
|
||||
models.Prefetch(
|
||||
query=[0.01, 0.45, 0.67, ...], # <-- dense vector
|
||||
query=[0.01, 0.45, 0.67], # <-- dense vector
|
||||
filter=models.Filter(
|
||||
must=models.FieldCondition(
|
||||
key="color",
|
||||
match=models.Match(value="green"),
|
||||
match=models.MatchValue(value="green"),
|
||||
),
|
||||
),
|
||||
limit=10,
|
||||
|
||||
@@ -41,7 +41,7 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="name_of_the_field_to_index",
|
||||
field_schema="keyword",
|
||||
field_schema=models.PayloadSchemaType.KEYWORD,
|
||||
)
|
||||
```
|
||||
|
||||
@@ -182,7 +182,7 @@ client.create_payload_index(
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient, Schemas } from "@qdrant/js-client-rest";
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
@@ -379,7 +379,7 @@ client.create_payload_index(
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient, Schemas } from "@qdrant/js-client-rest";
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
@@ -520,7 +520,7 @@ client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="payload_field_name",
|
||||
field_schema=models.KeywordIndexParams(
|
||||
type="keyword",
|
||||
type=models.KeywordIndexType.KEYWORD,
|
||||
on_disk=True,
|
||||
),
|
||||
)
|
||||
@@ -677,7 +677,7 @@ client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="payload_field_name",
|
||||
field_schema=models.KeywordIndexParams(
|
||||
type="keyword",
|
||||
type=models.KeywordIndexType.KEYWORD,
|
||||
is_tenant=True,
|
||||
),
|
||||
)
|
||||
@@ -814,8 +814,8 @@ PUT /collections/{collection_name}/index
|
||||
client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="timestamp",
|
||||
field_schema=models.KeywordIndexParams(
|
||||
type="integer",
|
||||
field_schema=models.IntegerIndexParams(
|
||||
type=models.IntegerIndexType.INTEGER,
|
||||
is_principal=True,
|
||||
),
|
||||
)
|
||||
@@ -834,7 +834,7 @@ client.createPayloadIndex("{collection_name}", {
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
CreateFieldIndexCollectionBuilder,
|
||||
IntegerdIndexParamsBuilder,
|
||||
IntegerIndexParamsBuilder,
|
||||
FieldType
|
||||
};
|
||||
use qdrant_client::{Qdrant, QdrantError};
|
||||
@@ -848,7 +848,7 @@ client.create_field_index(
|
||||
FieldType::Integer,
|
||||
)
|
||||
.field_index_params(
|
||||
IntegerdIndexParamsBuilder::default()
|
||||
IntegerIndexParamsBuilder::default()
|
||||
.is_principal(true),
|
||||
),
|
||||
);
|
||||
@@ -1009,12 +1009,11 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
sparse_vectors={
|
||||
"text": models.SparseVectorIndexParams(
|
||||
index=models.SparseVectorIndexType(
|
||||
on_disk=False,
|
||||
),
|
||||
),
|
||||
vectors_config={},
|
||||
sparse_vectors_config={
|
||||
"text": models.SparseVectorParams(
|
||||
index=models.SparseIndexParams(on_disk=False),
|
||||
)
|
||||
},
|
||||
)
|
||||
```
|
||||
@@ -1169,7 +1168,8 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
sparse_vectors={
|
||||
vectors_config={},
|
||||
sparse_vectors_config={
|
||||
"text": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
),
|
||||
|
||||
@@ -82,7 +82,7 @@ storage:
|
||||
# Memmap storage is disabled by default, to enable it, set this threshold to a reasonable value.
|
||||
# To disable memmap storage, set this to `0`.
|
||||
# Note: 1Kb = 1 vector of size 256
|
||||
memmap_threshold_kb: 200000
|
||||
memmap_threshold: 200000
|
||||
|
||||
# Maximum size (in kilobytes) of vectors allowed for plain index, exceeding this threshold will enable vector indexing
|
||||
# Default value is 20,000, based on <https://github.com/google-research/google-research/blob/master/scann/docs/algorithms.md>.
|
||||
|
||||
@@ -1386,7 +1386,7 @@ var client = new QdrantClient("localhost", 6334);
|
||||
await client.FacetAsync(
|
||||
"{collection_name}",
|
||||
key: "size",
|
||||
filter: MatchKeyword("color", "red"),
|
||||
filter: MatchKeyword("color", "red")
|
||||
);
|
||||
```
|
||||
|
||||
|
||||
@@ -1367,7 +1367,7 @@ client.delete_vectors(
|
||||
```typescript
|
||||
client.deleteVectors("{collection_name}", {
|
||||
points: [0, 3, 10],
|
||||
vectors: ["text", "image"],
|
||||
vector: ["text", "image"],
|
||||
});
|
||||
```
|
||||
|
||||
@@ -1705,13 +1705,6 @@ REST API ([Schema](https://api.qdrant.tech/api-reference/points/get-point)):
|
||||
GET /collections/{collection_name}/points/{point_id}
|
||||
```
|
||||
|
||||
<!--
|
||||
Python client:
|
||||
|
||||
```python
|
||||
```
|
||||
-->
|
||||
|
||||
## Scroll points
|
||||
|
||||
Sometimes it might be necessary to get all stored points without knowing ids, or iterate over points that correspond to a filter.
|
||||
@@ -1825,21 +1818,21 @@ import (
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client.Scroll(context.Background(), &qdrant.ScrollPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Filter: &qdrant.Filter{
|
||||
Must: []*qdrant.Condition{
|
||||
qdrant.NewMatch("color", "red"),
|
||||
},
|
||||
client.Scroll(context.Background(), &qdrant.ScrollPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Filter: &qdrant.Filter{
|
||||
Must: []*qdrant.Condition{
|
||||
qdrant.NewMatch("color", "red"),
|
||||
},
|
||||
Limit: qdrant.PtrOf(uint32(1)),
|
||||
WithPayload: qdrant.NewWithPayload(true),
|
||||
})
|
||||
},
|
||||
Limit: qdrant.PtrOf(uint32(1)),
|
||||
WithPayload: qdrant.NewWithPayload(true),
|
||||
})
|
||||
```
|
||||
|
||||
Returns all point with `color` = `red`.
|
||||
|
||||
@@ -1876,7 +1876,7 @@ from qdrant_client import QdrantClient, models
|
||||
|
||||
sampled = client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=models.SampleQuery(sample=models.Sample.Random)
|
||||
query=models.SampleQuery(sample=models.Sample.RANDOM)
|
||||
)
|
||||
```
|
||||
|
||||
|
||||
@@ -157,11 +157,11 @@ This will create a collection with all vectors immediately stored in memmap stor
|
||||
This is the recommended way, in case your Qdrant instance operates with fast disks and you are working with large collections.
|
||||
|
||||
|
||||
- Set up `memmap_threshold_kb` option (deprecated). This option will set the threshold after which the segment will be converted to memmap storage.
|
||||
- Set up `memmap_threshold` option. This option will set the threshold after which the segment will be converted to memmap storage.
|
||||
|
||||
There are two ways to do this:
|
||||
|
||||
1. You can set the threshold globally in the [configuration file](/documentation/guides/configuration/). The parameter is called `memmap_threshold_kb`.
|
||||
1. You can set the threshold globally in the [configuration file](/documentation/guides/configuration/). The parameter is called `memmap_threshold` (previously `memmap_threshold_kb`).
|
||||
2. You can set the threshold for each collection separately during [creation](/documentation/concepts/collections/#create-collection) or [update](/documentation/concepts/collections/#update-collection-parameters).
|
||||
|
||||
```http
|
||||
|
||||
@@ -108,6 +108,7 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config={},
|
||||
sparse_vectors_config={
|
||||
"text": models.SparseVectorParams(),
|
||||
},
|
||||
@@ -258,6 +259,7 @@ client.upsert("{collection_name}", {
|
||||
},
|
||||
},
|
||||
}
|
||||
]
|
||||
});
|
||||
```
|
||||
|
||||
@@ -322,8 +324,8 @@ await client.UpsertAsync(
|
||||
points: new List < PointStruct > {
|
||||
new() {
|
||||
Id = 1,
|
||||
Vectors = new Dictionary < string, Vector > {
|
||||
["text"] = ([0.1 f, 0.2 f, 0.3 f, 0.4 f], [1, 3, 5, 7])
|
||||
Vectors = new Dictionary <string, Vector> {
|
||||
["text"] = ([0.1f, 0.2f, 0.3f, 0.4f], [1, 3, 5, 7])
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -379,7 +381,7 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
result = client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query_vector=models.SparseVector(indices=[1, 3, 5, 7], values=[0.1, 0.2, 0.3, 0.4]),
|
||||
query=models.SparseVector(indices=[1, 3, 5, 7], values=[0.1, 0.2, 0.3, 0.4]),
|
||||
using="text",
|
||||
).points
|
||||
```
|
||||
@@ -534,7 +536,7 @@ client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(
|
||||
size=128,
|
||||
distance=models.Distance.Cosine,
|
||||
distance=models.Distance.COSINE,
|
||||
multivector_config=models.MultiVectorConfig(
|
||||
comparator=models.MultiVectorComparator.MAX_SIM
|
||||
),
|
||||
@@ -671,9 +673,9 @@ client.upsert(
|
||||
models.PointStruct(
|
||||
id=1,
|
||||
vector=[
|
||||
[-0.013, 0.020, -0.007, -0.111, ...],
|
||||
[-0.030, -0.055, 0.001, 0.072, ...],
|
||||
[-0.041, 0.014, -0.032, -0.062, ...]
|
||||
[-0.013, 0.020, -0.007, -0.111],
|
||||
[-0.030, -0.055, 0.001, 0.072],
|
||||
[-0.041, 0.014, -0.032, -0.062]
|
||||
],
|
||||
)
|
||||
],
|
||||
@@ -823,25 +825,24 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=[
|
||||
[-0.013, 0.020, -0.007, -0.111, ...],
|
||||
[-0.030, -0.055, 0.001, 0.072, ...],
|
||||
[-0.041, 0.014, -0.032, -0.062, ...]
|
||||
[-0.013, 0.020, -0.007, -0.111],
|
||||
[-0.030, -0.055, 0.001, 0.072],
|
||||
[-0.041, 0.014, -0.032, -0.062]
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
"query": [
|
||||
[-0.013, 0.020, -0.007, -0.111, ...],
|
||||
[-0.030, -0.055, 0.001, 0.072, ...],
|
||||
[-0.041, 0.014, -0.032, -0.062, ...]
|
||||
]
|
||||
"query": [
|
||||
[-0.013, 0.020, -0.007, -0.111],
|
||||
[-0.030, -0.055, 0.001, 0.072],
|
||||
[-0.041, 0.014, -0.032, -0.062]
|
||||
]
|
||||
});
|
||||
```
|
||||
|
||||
@@ -855,9 +856,9 @@ let res = client.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(VectorInput::new_multi(
|
||||
vec![
|
||||
vec![-0.013, 0.020, -0.007, -0.111, ...],
|
||||
vec![-0.030, -0.055, 0.001, 0.072, ...],
|
||||
vec![-0.041, 0.014, -0.032, -0.062, ...],
|
||||
vec![-0.013, 0.020, -0.007, -0.111],
|
||||
vec![-0.030, -0.055, 0.001, 0.072],
|
||||
vec![-0.041, 0.014, -0.032, -0.062],
|
||||
]
|
||||
))
|
||||
).await?;
|
||||
@@ -924,7 +925,7 @@ client.Query(context.Background(), &qdrant.QueryPoints{
|
||||
|
||||
## Named Vectors
|
||||
|
||||
In Qdrant, you can store multiple vectors of different sizes in the same data [point](/documentation/concepts/points/). This is useful when you need to define your data with multiple embeddings to represent different features or modalities (e.g., image, text or video).
|
||||
In Qdrant, you can store multiple vectors of different sizes and [types](#vector-types) in the same data [point](/documentation/concepts/points/). This is useful when you need to define your data with multiple embeddings to represent different features or modalities (e.g., image, text or video).
|
||||
|
||||
To store different vectors for each point, you need to create separate named vector spaces in the [collection](/documentation/concepts/collections/). You can define these vector spaces during collection creation and manage them independently.
|
||||
|
||||
@@ -937,48 +938,34 @@ To create a collection with named vectors, you need to specify a configuration f
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"image": {
|
||||
"size": 4,
|
||||
"distance": "Dot"
|
||||
},
|
||||
"text": {
|
||||
"size": 8,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
"vectors": {
|
||||
"image": {
|
||||
"size": 4,
|
||||
"distance": "Dot"
|
||||
},
|
||||
"text": {
|
||||
"size": 5,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
},
|
||||
"sparse_vectors": {
|
||||
"text-sparse": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```bash
|
||||
curl -X PUT http://localhost:6333/collections/{collection_name} \
|
||||
-H 'Content-Type: application/json' \
|
||||
--data-raw '{
|
||||
"vectors": {
|
||||
"image": {
|
||||
"size": 4,
|
||||
"distance": "Dot"
|
||||
},
|
||||
"text": {
|
||||
"size": 8,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config={
|
||||
"image": models.VectorParams(size=4, distance=models.Distance.DOT),
|
||||
"text": models.VectorParams(size=8, distance=models.Distance.COSINE),
|
||||
"text": models.VectorParams(size=5, distance=models.Distance.COSINE),
|
||||
},
|
||||
sparse_vectors_config={"text-sparse": models.SparseVectorParams()},
|
||||
)
|
||||
```
|
||||
|
||||
@@ -988,28 +975,38 @@ import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
vectors: {
|
||||
image: { size: 4, distance: "Dot" },
|
||||
text: { size: 8, distance: "Cosine" },
|
||||
},
|
||||
vectors: {
|
||||
image: { size: 4, distance: "Dot" },
|
||||
text: { size: 5, distance: "Cosine" },
|
||||
},
|
||||
sparse_vectors: {
|
||||
text_sparse: {}
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
CreateCollectionBuilder, Distance, VectorParamsBuilder, VectorsConfigBuilder,
|
||||
CreateCollectionBuilder, Distance, SparseVectorParamsBuilder, SparseVectorsConfigBuilder,
|
||||
VectorParamsBuilder, VectorsConfigBuilder,
|
||||
};
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
let mut vector_config = VectorsConfigBuilder::default();
|
||||
vector_config.add_named_vector_params("text", VectorParamsBuilder::new(4, Distance::Dot));
|
||||
vector_config.add_named_vector_params("image", VectorParamsBuilder::new(8, Distance::Cosine));
|
||||
vector_config.add_named_vector_params("text", VectorParamsBuilder::new(5, Distance::Dot));
|
||||
vector_config.add_named_vector_params("image", VectorParamsBuilder::new(4, Distance::Cosine));
|
||||
|
||||
let mut sparse_vectors_config = SparseVectorsConfigBuilder::default();
|
||||
sparse_vectors_config
|
||||
.add_named_vector_params("text-sparse", SparseVectorParamsBuilder::default());
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}").vectors_config(vector_config),
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.vectors_config(vector_config)
|
||||
.sparse_vectors_config(sparse_vectors_config),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
@@ -1019,19 +1016,35 @@ import java.util.Map;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Distance;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorConfig;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorParamsMap;
|
||||
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
"{collection_name}",
|
||||
Map.of(
|
||||
"image", VectorParams.newBuilder().setSize(4).setDistance(Distance.Dot).build(),
|
||||
"text",
|
||||
VectorParams.newBuilder().setSize(8).setDistance(Distance.Cosine).build()))
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setVectorsConfig(VectorsConfig.newBuilder().setParamsMap(
|
||||
VectorParamsMap.newBuilder().putAllMap(Map.of("image",
|
||||
VectorParams.newBuilder()
|
||||
.setSize(4)
|
||||
.setDistance(Distance.Dot)
|
||||
.build(),
|
||||
"text",
|
||||
VectorParams.newBuilder()
|
||||
.setSize(5)
|
||||
.setDistance(Distance.Cosine)
|
||||
.build()))))
|
||||
.setSparseVectorsConfig(SparseVectorConfig.newBuilder().putMap(
|
||||
"text-sparse", SparseVectorParams.getDefaultInstance()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
@@ -1043,15 +1056,22 @@ var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
vectorsConfig: new VectorParamsMap {
|
||||
Map = {
|
||||
vectorsConfig: new VectorParamsMap
|
||||
{
|
||||
Map = {
|
||||
["image"] = new VectorParams {
|
||||
Size = 4, Distance = Distance.Dot
|
||||
},
|
||||
["text"] = new VectorParams {
|
||||
Size = 8, Distance = Distance.Cosine
|
||||
Size = 5, Distance = Distance.Cosine
|
||||
},
|
||||
}
|
||||
},
|
||||
sparseVectorsConfig: new SparseVectorConfig
|
||||
{
|
||||
Map = {
|
||||
["text-sparse"] = new()
|
||||
}
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -1077,24 +1097,33 @@ client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||
Distance: qdrant.Distance_Dot,
|
||||
},
|
||||
"text": {
|
||||
Size: 8,
|
||||
Size: 5,
|
||||
Distance: qdrant.Distance_Cosine,
|
||||
},
|
||||
}),
|
||||
SparseVectorsConfig: qdrant.NewSparseVectorsConfig(
|
||||
map[string]*qdrant.SparseVectorParams{
|
||||
"text-sparse": {},
|
||||
},
|
||||
),
|
||||
})
|
||||
```
|
||||
|
||||
To insert a point with named vectors:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}/points
|
||||
PUT /collections/{collection_name}/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"image": [0.9, 0.1, 0.1, 0.2],
|
||||
"text": [0.4, 0.7, 0.1, 0.8, 0.1, 0.1, 0.9, 0.2]
|
||||
"text": [0.4, 0.7, 0.1, 0.8, 0.1],
|
||||
"text-sparse": {
|
||||
"indices": [1, 3, 5, 7],
|
||||
"values": [0.1, 0.2, 0.3, 0.4]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
@@ -1109,7 +1138,11 @@ client.upsert(
|
||||
id=1,
|
||||
vector={
|
||||
"image": [0.9, 0.1, 0.1, 0.2],
|
||||
"text": [0.4, 0.7, 0.1, 0.8, 0.1, 0.1, 0.9, 0.2],
|
||||
"text": [0.4, 0.7, 0.1, 0.8, 0.1],
|
||||
"text-sparse": {
|
||||
"indices": [1, 3, 5, 7],
|
||||
"values": [0.1, 0.2, 0.3, 0.4],
|
||||
},
|
||||
},
|
||||
),
|
||||
],
|
||||
@@ -1118,41 +1151,44 @@ client.upsert(
|
||||
|
||||
```typescript
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
image: [0.9, 0.1, 0.1, 0.2],
|
||||
text: [0.4, 0.7, 0.1, 0.8, 0.1, 0.1, 0.9, 0.2],
|
||||
},
|
||||
},
|
||||
],
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
image: [0.9, 0.1, 0.1, 0.2],
|
||||
text: [0.4, 0.7, 0.1, 0.8, 0.1],
|
||||
text_sparse: {
|
||||
indices: [1, 3, 5, 7],
|
||||
values: [0.1, 0.2, 0.3, 0.4]
|
||||
}
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use std::collections::HashMap;
|
||||
|
||||
use qdrant_client::qdrant::{PointStruct, UpsertPointsBuilder};
|
||||
use qdrant_client::qdrant::{
|
||||
NamedVectors, PointStruct, UpsertPointsBuilder, Vector,
|
||||
};
|
||||
use qdrant_client::Payload;
|
||||
|
||||
client
|
||||
.upsert_points(
|
||||
UpsertPointsBuilder::new(
|
||||
"{collection_name}",
|
||||
vec![
|
||||
PointStruct::new(
|
||||
1,
|
||||
HashMap::from([
|
||||
("image".to_string(), vec![0.9, 0.1, 0.1, 0.2]),
|
||||
(
|
||||
"text".to_string(),
|
||||
vec![0.4, 0.7, 0.1, 0.8, 0.1, 0.1, 0.9, 0.2],
|
||||
),
|
||||
]),
|
||||
Payload::default(),
|
||||
),
|
||||
],
|
||||
vec![PointStruct::new(
|
||||
1,
|
||||
NamedVectors::default()
|
||||
.add_vector("text", Vector::new_dense(vec![0.4, 0.7, 0.1, 0.8, 0.1]))
|
||||
.add_vector("image", Vector::new_dense(vec![0.9, 0.1, 0.1, 0.2]))
|
||||
.add_vector(
|
||||
"text-sparse",
|
||||
Vector::new_sparse(vec![1, 3, 5, 7], vec![0.1, 0.2, 0.3, 0.4]),
|
||||
),
|
||||
Payload::default(),
|
||||
)],
|
||||
)
|
||||
.wait(true),
|
||||
)
|
||||
@@ -1181,7 +1217,9 @@ client
|
||||
"image",
|
||||
vector(List.of(0.9f, 0.1f, 0.1f, 0.2f)),
|
||||
"text",
|
||||
vector(List.of(0.4f, 0.7f, 0.1f, 0.8f, 0.1f, 0.1f, 0.9f, 0.2f)))))
|
||||
vector(List.of(0.4f, 0.7f, 0.1f, 0.8f, 0.1f)),
|
||||
"text-sparse",
|
||||
vector(List.of(0.1f, 0.2f, 0.3f, 0.4f), List.of(1, 3, 5, 7)))))
|
||||
.build()))
|
||||
.get();
|
||||
```
|
||||
@@ -1190,22 +1228,25 @@ client
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.UpsertAsync(
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Dictionary<string, float[]>
|
||||
{
|
||||
["image"] = [0.9f, 0.1f, 0.1f, 0.2f],
|
||||
["text"] = [0.4f, 0.7f, 0.1f, 0.8f, 0.1f, 0.1f, 0.9f, 0.2f]
|
||||
}
|
||||
}
|
||||
}
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Dictionary<string, Vector>
|
||||
{
|
||||
["image"] = new() {
|
||||
Data = {0.9f, 0.1f, 0.1f, 0.2f}
|
||||
},
|
||||
["text"] = new() {
|
||||
Data = {0.4f, 0.7f, 0.1f, 0.8f, 0.1f}
|
||||
},
|
||||
["text-sparse"] = ([0.1f, 0.2f, 0.3f, 0.4f], [1, 3, 5, 7]),
|
||||
}
|
||||
}
|
||||
}
|
||||
);
|
||||
```
|
||||
|
||||
@@ -1216,11 +1257,6 @@ import (
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client.Upsert(context.Background(), &qdrant.UpsertPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Points: []*qdrant.PointStruct{
|
||||
@@ -1228,13 +1264,17 @@ client.Upsert(context.Background(), &qdrant.UpsertPoints{
|
||||
Id: qdrant.NewIDNum(1),
|
||||
Vectors: qdrant.NewVectorsMap(map[string]*qdrant.Vector{
|
||||
"image": qdrant.NewVector(0.9, 0.1, 0.1, 0.2),
|
||||
"text": qdrant.NewVector(0.4, 0.7, 0.1, 0.8, 0.1, 0.1, 0.9, 0.2),
|
||||
"text": qdrant.NewVector(0.4, 0.7, 0.1, 0.8, 0.1),
|
||||
"text-sparse": qdrant.NewVectorSparse(
|
||||
[]uint32{1, 3, 5, 7},
|
||||
[]float32{0.1, 0.2, 0.3, 0.4}),
|
||||
}),
|
||||
},
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
You can also upload
|
||||
To search with named vectors (available in `query` API):
|
||||
|
||||
```http
|
||||
@@ -1598,13 +1638,11 @@ client = QdrantClient(url="http://localhost:6333")
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(
|
||||
size=128,
|
||||
distance=models.Distance.COSINE,
|
||||
datatype=models.Datatype.UINT8
|
||||
size=128, distance=models.Distance.COSINE, datatype=models.Datatype.UINT8
|
||||
),
|
||||
sparse_vectors_config={
|
||||
"text": models.SparseVectorParams(
|
||||
index=models.SparseIndexConfig(datatype=models.Datatype.UINT8)
|
||||
index=models.SparseIndexParams(datatype=models.Datatype.UINT8)
|
||||
),
|
||||
},
|
||||
)
|
||||
|
||||
@@ -0,0 +1,373 @@
|
||||
---
|
||||
title: Data Ingestion for Beginners
|
||||
weight: 11
|
||||
partition: build
|
||||
social_preview_image: /documentation/examples/data-ingestion-beginners/social_preview.png
|
||||
---
|
||||

|
||||
|
||||
# Send S3 Data to Qdrant Vector Store with LangChain
|
||||
|
||||
| Time: 30 min | Level: Beginner | | |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
**Data ingestion into a vector store** is essential for building effective search and retrieval algorithms, especially since nearly 80% of data is unstructured, lacking any predefined format.
|
||||
|
||||
In this tutorial, we’ll create a streamlined data ingestion pipeline, pulling data directly from **AWS S3** and feeding it into Qdrant. We’ll dive into vector embeddings, transforming unstructured data into a format that allows you to search documents semantically. Prepare to discover new ways to uncover insights hidden within unstructured data!
|
||||
|
||||
## Ingestion Workflow Architecture
|
||||
|
||||
We’ll set up a powerful document ingestion and analysis pipeline in this workflow using cloud storage, natural language processing (NLP) tools, and embedding technologies. Starting with raw data in an S3 bucket, we'll preprocess it with LangChain, apply embedding APIs for both text and images and store the results in Qdrant – a vector database optimized for similarity search.
|
||||
|
||||
**Figure 1: Data Ingestion Workflow Architecture**
|
||||
|
||||

|
||||
|
||||
Let's break down each component of this workflow:
|
||||
|
||||
- **S3 Bucket:** This is our starting point—a centralized, scalable storage solution for various file types like PDFs, images, and text.
|
||||
- **LangChain:** Acting as the pipeline’s orchestrator, LangChain handles extraction, preprocessing, and manages data flow for embedding generation. It simplifies processing PDFs, so you won’t need to worry about applying OCR (Optical Character Recognition) here.
|
||||
- **Text Embeddings API:** This API transforms text from files and PDFs into vector representations, capturing semantic meaning for advanced analysis and search.
|
||||
- **Image Embeddings API:** This tool takes image files and converts them into vector representations, making similarity search and analysis possible for visual content.
|
||||
- **Qdrant:** As your vector database, Qdrant stores embeddings and their [payloads](https://qdrant.tech/documentation/concepts/payload/), enabling efficient similarity search and retrieval across all content types.
|
||||
|
||||
## Prerequisites
|
||||

|
||||
|
||||
In this section, you’ll get a step-by-step guide on ingesting data from an S3 bucket. But before we dive in, let’s make sure you’re set up with all the prerequisites:
|
||||
|
||||
|||
|
||||
|-|-|
|
||||
|Sample Data|We’ll use a sample dataset, where each folder includes product reviews in text format along with corresponding images.|
|
||||
|AWS Account|An active [AWS account](https://aws.amazon.com/free/) with access to S3 services.|
|
||||
|Qdrant Account|A [Qdrant Cloud account](https://cloud.qdrant.io) with access to the WebUI for managing collections and running queries.|
|
||||
|Embedding Models|Use embedding models such as [OpenAI's text embedding model](https://openai.com/index/openai-api/) for text files and CLIP for images.|
|
||||
|LangChain| You will use this [popular framework](https://www.langchain.com) to tie everything together. |
|
||||
|
||||
#### Supported Document Types
|
||||
|
||||
The documents used for ingestion can be of various types, such as PDFs, text files, or images. We will organize a structured S3 bucket with folders with the supported document types for testing and experimentation.
|
||||
|
||||
#### Python Environment
|
||||
|
||||
Ensure you have a Python environment (Python 3.9 or higher) with these libraries installed:
|
||||
|
||||
```jsx
|
||||
boto3
|
||||
langchain-community
|
||||
langchain
|
||||
python-dotenv
|
||||
unstructured
|
||||
unstructured[pdf]
|
||||
qdrant_client
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
**Access Keys:** Store your AWS access key, S3 secret key, and Qdrant API key in a .env file for easy access. You’ll also need an **OpenAI API key**. Here’s a sample `.env` file.
|
||||
|
||||
```text
|
||||
ACCESS_KEY = ""
|
||||
SECRET_ACCESS_KEY = ""
|
||||
OPENAI_API_KEY = ""
|
||||
QDRANT_KEY = ""
|
||||
```
|
||||
---
|
||||
|
||||
<aside role="alert"> Although the code includes support for processing PDFs, the sample data currently has no PDF files included. </aside>
|
||||
|
||||
## Step 1: Ingesting Data from S3
|
||||
|
||||

|
||||
|
||||
The LangChain framework makes it easy to ingest data from storage services like AWS S3, with built-in support for loading documents in formats such as PDFs, images, and text files.
|
||||
|
||||
To connect LangChain with S3, you’ll use the `S3DirectoryLoader`, which lets you load files directly from an S3 bucket into LangChain’s pipeline.
|
||||
|
||||
### Example: Configuring LangChain to Load Files from S3
|
||||
|
||||
Here’s how to set up LangChain to ingest data from an S3 bucket:
|
||||
|
||||
```jsx
|
||||
from langchain_community.document_loaders import S3DirectoryLoader
|
||||
|
||||
# Initialize the S3 document loader
|
||||
loader = S3DirectoryLoader(
|
||||
"product-dataset", # S3 bucket name
|
||||
"p_1", #S3 Folder name containing the data for the first product
|
||||
aws_access_key_id=aws_access_key_id, # AWS Access Key
|
||||
aws_secret_access_key=aws_secret_access_key # AWS Secret Access Key
|
||||
)
|
||||
|
||||
# Load documents from the specified S3 bucket
|
||||
docs = loader.load()
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 2. Turning Documents into Embeddings
|
||||
|
||||
[Embeddings](/articles/what-are-embeddings/) are the secret sauce here—they’re numerical representations of data (like text, images, or audio) that capture the “meaning” in a form that’s easy to compare. By converting text and images into embeddings, you’ll be able to perform similarity searches quickly and efficiently. Think of embeddings as the bridge to storing and retrieving meaningful insights from your data in Qdrant.
|
||||
|
||||
### Models We’ll Use for Generating Embeddings
|
||||
|
||||
To get things rolling, we’ll use two powerful models:
|
||||
|
||||
1. **OpenAI Embeddings** for transforming text data.
|
||||
2. **CLIP (Contrastive Language-Image Pretraining)** for image data.
|
||||
|
||||
### Text Embeddings
|
||||
|
||||
Here, we’re using `OpenAIEmbeddings`—a pre-trained model that turns text into embeddings, capturing its underlying meaning. With these, we’re ready to unlock powerful search and retrieval tasks.
|
||||
|
||||
Here’s how to set up the OpenAI embedding model for text:
|
||||
|
||||
```jsx
|
||||
from langchain_openai import OpenAIEmbeddings
|
||||
|
||||
# Initialize the text embedding model from OpenAI
|
||||
text_embedding_model = OpenAIEmbeddings(model="text-embedding-3-small")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
The `text-embedding-3-small` model generates high-quality text embeddings, making it easier to find documents with similar meanings—even if they don’t contain the exact same words.
|
||||
|
||||
### Image Embeddings
|
||||
|
||||
We’re using OpenAI’s `ClipModel`, which is designed to handle both images and text. Here, we’ll use CLIP to generate embeddings for images, allowing you to compare image content based on its semantic meaning.
|
||||
|
||||
Here’s how to set up the CLIP model and processor::
|
||||
|
||||
```jsx
|
||||
From transformers import CLIPProcessor, CLIPModel
|
||||
import torch
|
||||
|
||||
# Initialize the CLIP model and processor
|
||||
clip_model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
|
||||
clip_processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
The `clip-vit-base-patch32` model is specifically trained to align images and text within the same embedding space, meaning that similar images will have embeddings that are close to each other. With the models set up, you can now create a function to process documents based on their type and generate embeddings using the CLIP model, for instance, to convert image features into vectors.
|
||||
|
||||
```jsx
|
||||
def embed_image_with_clip(image):
|
||||
inputs = clip_processor(images=image, return_tensors="pt")
|
||||
with torch.no_grad():
|
||||
image_features = clip_model.get_image_features(**inputs)
|
||||
return image_features.cpu().numpy()
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Document Processing Function
|
||||
|
||||

|
||||
|
||||
Next, we’ll create a `process_document` function that converts various file types—such as text, PDF, and image—into embeddings. This function applies different methods to extract and embed content based on the file type, making it versatile for handling multiple formats effectively.
|
||||
|
||||
```jsx
|
||||
def process_document(doc):
|
||||
source = doc.metadata['source'] # Extract document source (e.g., S3 URL)
|
||||
|
||||
# Processing Text Files
|
||||
if source.endswith('.txt'):
|
||||
text = doc.page_content # Extract the content from the text file
|
||||
print(f"Processing .txt file: {source}")
|
||||
return text, text_embedding_model.embed_documents([text]) # Convert to embeddings
|
||||
|
||||
# Processing PDF Files
|
||||
elif source.endswith('.pdf'):
|
||||
content = doc.page_content # Extract content from the PDF
|
||||
print(f"Processing .pdf file: {source}")
|
||||
return content, text_embedding_model.embed_documents([content]) # Convert to embeddings
|
||||
|
||||
# Processing Image Files
|
||||
elif source.endswith('.png'):
|
||||
print(f"Processing .png file: {source}")
|
||||
bucket_name, object_key = parse_s3_url(source) # Parse the S3 URL
|
||||
response = s3.get_object(Bucket=bucket_name, Key=object_key) # Fetch image from S3
|
||||
img_bytes = response['Body'].read()
|
||||
|
||||
# Load the image and convert to embeddings
|
||||
img = Image.open(io.BytesIO(img_bytes))
|
||||
return source, embed_image_with_clip(img) # Convert to image embeddings
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
Here’s the explanation of the code above:
|
||||
|
||||
- **File Type Detection**
|
||||
|
||||
First, the function checks the file type using the file extension (.txt, .pdf, .png) from the document's source metadata. This tells it how to process the content and which embedding model to use
|
||||
|
||||
- **File Processing**
|
||||
- **Text Files (.txt):** For text files, it’s straightforward—the content is extracted as plain text and passed to the OpenAIEmbeddings model. Then, the embed_documents() function transforms this text into numerical vectors, capturing its meaning in embedding form.
|
||||
- **PDFs:** PDFs often have a lot to offer, like multi-page content. Here, we’re using LangChain’s document loader to extract the text from each page, and then converting it into embeddings with the OpenAI text embedding model. This lets you capture the full richness of the document.
|
||||
- **Images (.png):** For images (e.g., .png files), the function fetches the image from S3 using its URL. Once loaded with the Python Imaging Library (PIL), the image is passed to the CLIP model, which generates an embedding that represents the image’s semantic features.
|
||||
|
||||
### Helper Functions for Document Processing
|
||||
|
||||
To retrieve images from S3, a helper function `parse_s3_url` breaks down the S3 URL into its bucket and critical components. This is essential for fetching the image from S3 storage.
|
||||
|
||||
```jsx
|
||||
def parse_s3_url(s3_url):
|
||||
parts = s3_url.replace("s3://", "").split("/", 1)
|
||||
bucket_name = parts[0]
|
||||
object_key = parts[1]
|
||||
return bucket_name, object_key
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 3: Loading Embeddings into Qdrant
|
||||
|
||||

|
||||
|
||||
Now that your documents have been processed and converted into embeddings, the next step is to load these embeddings into Qdrant.
|
||||
|
||||
### Creating a Collection in Qdrant
|
||||
|
||||
In Qdrant, data is organized in collections, each representing a set of embeddings (or points) and their associated metadata (payload). To store the embeddings generated earlier, you’ll first need to create a collection.
|
||||
|
||||
Here’s how to create a collection in Qdrant to store both text and image embeddings:
|
||||
|
||||
```jsx
|
||||
def create_collection(collection_name):
|
||||
qdrant_client.create_collection(
|
||||
collection_name,
|
||||
vectors_config={
|
||||
"text_embedding": models.VectorParams(
|
||||
size=1536, # Dimension of text embeddings
|
||||
distance=models.Distance.COSINE, # Cosine similarity is used for comparison
|
||||
),
|
||||
"image_embedding": models.VectorParams(
|
||||
size=512, # Dimension of image embeddings
|
||||
distance=models.Distance.COSINE, # Cosine similarity is used for comparison
|
||||
),
|
||||
},
|
||||
)
|
||||
|
||||
create_collection("products-data")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
This function creates a collection for storing text (1536 dimensions) and image (512 dimensions) embeddings, using cosine similarity to compare embeddings within the collection.
|
||||
|
||||
Once the collection is set up, you can load the embeddings into Qdrant. This involves inserting (or updating) the embeddings and their associated metadata (payload) into the specified collection.
|
||||
|
||||
Here’s the code for loading embeddings into Qdrant:
|
||||
|
||||
```jsx
|
||||
def ingest_data(points):
|
||||
operation_info = qdrant_client.upsert(
|
||||
collection_name="products-data", # Collection where data is being inserted
|
||||
points=points
|
||||
)
|
||||
return operation_info
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
**Explanation of Ingestion**
|
||||
|
||||
1. **Upserting the Data Point:** The upsert method on the `qdrant_client` inserts each PointStruct into the specified collection. If a point with the same ID already exists, it will be updated with the new values.
|
||||
2. **Operation Info:** The function returns `operation_info`, which contains details about the upsert operation, such as success status or any potential errors.
|
||||
|
||||
**Running the Ingestion Code**
|
||||
|
||||
Here’s how to call the function and ingest data:
|
||||
|
||||
```jsx
|
||||
if __name__ == "__main__":
|
||||
collection_name = "products-data"
|
||||
create_collection(collection_name)
|
||||
for i in range(1,6): # Five documents
|
||||
folder = f"p_{i}"
|
||||
loader = S3DirectoryLoader(
|
||||
"product-dataset",
|
||||
folder,
|
||||
aws_access_key_id=aws_access_key_id,
|
||||
aws_secret_access_key=aws_secret_access_key
|
||||
)
|
||||
docs = loader.load()
|
||||
text_embedding, image_embedding, points, text_review, product_image = [], [], [], "", ""
|
||||
for idx, doc in enumerate(docs):
|
||||
source = doc.metadata['source']
|
||||
if source.endswith(".txt"):
|
||||
text_review, text_embedding = process_document(doc)
|
||||
elif source.endswith(".png"):
|
||||
product_image, image_embedding = process_document(doc)
|
||||
if text_review:
|
||||
point = PointStruct(
|
||||
id=idx, # Unique identifier for each point
|
||||
vector={
|
||||
"text_embedding": text_embedding[0],
|
||||
"image_embedding": image_embedding[0].tolist(),
|
||||
},
|
||||
payload={
|
||||
"review": text_review,
|
||||
"product_image": product_image
|
||||
}
|
||||
)
|
||||
points.append(point)
|
||||
operation_info = ingest_data(points)
|
||||
print(operation_info)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
The `PointStruct` is instantiated with these key parameters:
|
||||
|
||||
- **id:** A unique identifier for each embedding, typically an incremental index.
|
||||
- **vector:** A dictionary containing the text and image embeddings generated from each document.
|
||||
- **payload:** A dictionary storing additional metadata, like product reviews and image references, which is invaluable for retrieval and context during searches.
|
||||
|
||||
The code dynamically loads folders from an S3 bucket, processes text and image files separately, and stores their embeddings and associated data in dedicated lists. It then creates a `PointStruct` for each data entry and calls the ingestion function to load it into Qdrant.
|
||||
|
||||
### Exploring the Qdrant WebUI Dashboard
|
||||
|
||||
Once the embeddings are loaded into Qdrant, you can use the WebUI dashboard to visualize and manage your collections. The dashboard provides a clear, structured interface for viewing collections and their data. Let’s take a closer look in the next section.
|
||||
|
||||
## Step 4: Visualizing Data in Qdrant WebUI
|
||||
|
||||
To start visualizing your data in the Qdrant WebUI, head to the **Overview** section and select **Access the database**.
|
||||
|
||||
**Figure 2: Accessing the Database from the Qdrant UI**
|
||||

|
||||
|
||||
When prompted, enter your API key. Once inside, you’ll be able to view your collections and the corresponding data points. You should see your collection displayed like this:
|
||||
|
||||
**Figure 3: The product-data Collection in Qdrant**
|
||||

|
||||
|
||||
Here’s a look at the most recent point ingested into Qdrant:
|
||||
|
||||
**Figure 4: The Latest Point Added to the product-data Collection**
|
||||

|
||||
|
||||
The Qdrant WebUI’s search functionality allows you to perform vector searches across your collections. With options to apply filters and parameters, retrieving relevant embeddings and exploring relationships within your data becomes easy. To start, head over to the **Console** in the left panel, where you can create queries:
|
||||
|
||||
**Figure 5: Overview of Console in Qdrant**
|
||||

|
||||
|
||||
The first query retrieves all collections, the second fetches points from the product-data collection, and the third performs a sample query. This demonstrates how straightforward it is to interact with your data in the Qdrant UI.
|
||||
|
||||
Now, let’s retrieve some documents from the database using a query!.
|
||||
|
||||
**Figure 6: Querying the Qdrant Client to Retrieve Relevant Documents**
|
||||

|
||||
|
||||
In this example, we queried **Phones with improved design**. Then, we converted the text to vectors using OpenAI and retrieved a relevant phone review highlighting design improvements.
|
||||
|
||||
## Conclusion
|
||||
|
||||
In this guide, we set up an S3 bucket, ingested various data types, and stored embeddings in Qdrant. Using LangChain, we dynamically processed text and image files, making it easy to work with each file type.
|
||||
|
||||
Now, it’s your turn. Try experimenting with different data types, such as videos, and explore Qdrant’s advanced features to enhance your applications. To get started, [sign up](https://cloud.qdrant.io/) for Qdrant today.
|
||||
|
||||

|
||||
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: Data Management
|
||||
weight: 18
|
||||
partition: build
|
||||
---
|
||||
|
||||
## Data Management Integrations
|
||||
|
||||
@@ -13,27 +13,25 @@ You can set up the Qdrant-Spark Connector in a few different ways, depending on
|
||||
|
||||
### GitHub Releases
|
||||
|
||||
The simplest way to get started is by downloading pre-packaged JAR file releases from the [GitHub releases page](https://github.com/qdrant/qdrant-spark/releases). These JAR files come with all the necessary dependencies.
|
||||
You can download the packaged JAR file from the [GitHub releases](https://github.com/qdrant/qdrant-spark/releases). It comes with all the required dependencies.
|
||||
|
||||
### Building from Source
|
||||
|
||||
If you prefer to build the JAR from source, you'll need [JDK 8](https://www.azul.com/downloads/#zulu) and [Maven](https://maven.apache.org/) installed on your system. Once you have the prerequisites in place, navigate to the project's root directory and run the following command:
|
||||
To build the JAR from source, you'll need [JDK 8](https://www.azul.com/downloads/#zulu) and [Maven](https://maven.apache.org/) installed on your system. Once you have those in place, navigate to the project's root directory and run the following command:
|
||||
|
||||
```bash
|
||||
mvn package
|
||||
mvn package -DskipTests
|
||||
```
|
||||
|
||||
This command will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default.
|
||||
This will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default.
|
||||
|
||||
### Maven Central
|
||||
|
||||
For use with Java and Scala projects, the package can be found [here](https://central.sonatype.com/artifact/io.qdrant/spark).
|
||||
The package can be found [here](https://central.sonatype.com/artifact/io.qdrant/spark).
|
||||
|
||||
## Usage
|
||||
|
||||
Below, we'll walk through the steps of creating a Spark session with Qdrant support and loading data into Qdrant.
|
||||
|
||||
### Creating a single-node Spark session with Qdrant Support
|
||||
Below, we'll walk through the steps of creating a Spark session and ingesting data into Qdrant.
|
||||
|
||||
To begin, import the necessary libraries and create a Spark session with Qdrant support:
|
||||
|
||||
@@ -42,7 +40,7 @@ from pyspark.sql import SparkSession
|
||||
|
||||
spark = SparkSession.builder.config(
|
||||
"spark.jars",
|
||||
"spark-VERSION.jar", # Specify the downloaded JAR file
|
||||
"spark-VERSION.jar", # Specify the path to the downloaded JAR file
|
||||
)
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
@@ -53,7 +51,7 @@ spark = SparkSession.builder.config(
|
||||
import org.apache.spark.sql.SparkSession
|
||||
|
||||
val spark = SparkSession.builder
|
||||
.config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file
|
||||
.config("spark.jars", "spark-VERSION.jar") // Specify the path to the downloaded JAR file
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate()
|
||||
@@ -65,7 +63,7 @@ import org.apache.spark.sql.SparkSession;
|
||||
public class QdrantSparkJavaExample {
|
||||
public static void main(String[] args) {
|
||||
SparkSession spark = SparkSession.builder()
|
||||
.config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file
|
||||
.config("spark.jars", "spark-VERSION.jar") // Specify the path to the downloaded JAR file
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate();
|
||||
@@ -73,8 +71,6 @@ public class QdrantSparkJavaExample {
|
||||
}
|
||||
```
|
||||
|
||||
### Loading data into Qdrant
|
||||
|
||||
<aside role="status">Before loading the data using this connector, a collection has to be <a href="/documentation/concepts/collections/#create-a-collection">created</a> in advance with the appropriate vector dimensions and configurations.</aside>
|
||||
|
||||
The connector supports ingesting multiple named/unnamed, dense/sparse vectors.
|
||||
@@ -229,7 +225,7 @@ You can use the `qdrant-spark` connector as a library in [Databricks](https://ww
|
||||
|
||||
## Datatype Support
|
||||
|
||||
Qdrant supports all the Spark data types, and the appropriate data types are mapped based on the provided schema.
|
||||
Qdrant supports most Spark data types, and the appropriate data types are mapped based on the provided schema.
|
||||
|
||||
## Configuration Options
|
||||
|
||||
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
title: Using the Database
|
||||
weight: 18
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Database Tutorials
|
||||
|
||||
| |
|
||||
|--------------------------------------------|
|
||||
| [Bulk Upload Vectors to a Qdrant Collection](/documentation/database-tutorials/bulk-upload/) |
|
||||
| [Backup and Restore Qdrant Collections Using Snapshots](/documentation/database-tutorials/create-snapshot/) |
|
||||
| [Load and Search Hugging Face Datasets with Qdrant](/documentation/database-tutorials/huggingface-datasets/) |
|
||||
| [Using Qdrant’s Async API for Efficient Python Applications](/documentation/database-tutorials/async-api/) |
|
||||
+5
-3
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Asynchronous API
|
||||
weight: 14
|
||||
title: Build With Async API
|
||||
aliases:
|
||||
- /documentation/tutorials/async-api/
|
||||
weight: 4
|
||||
---
|
||||
|
||||
# Using Qdrant asynchronously
|
||||
# Using Qdrant’s Async API for Efficient Python Applications
|
||||
|
||||
Asynchronous programming is being broadly adopted in the Python ecosystem. Tools such as FastAPI [have embraced this new
|
||||
paradigm](https://fastapi.tiangolo.com/async/), but it is also becoming a standard for ML models served as SaaS. For example, the Cohere SDK
|
||||
+5
-3
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Bulk Upload Vectors
|
||||
weight: 13
|
||||
aliases:
|
||||
- /documentation/tutorials/bulk-upload/
|
||||
weight: 1
|
||||
---
|
||||
|
||||
# Bulk upload a large number of vectors
|
||||
# Bulk Upload Vectors to a Qdrant Collection
|
||||
|
||||
Uploading a large-scale dataset fast might be a challenge, but Qdrant has a few tricks to help you with that.
|
||||
|
||||
@@ -110,7 +112,7 @@ memmaps may be enabled on a per-vector basis using the `on_disk` parameter. This
|
||||
will store vector data directly on disk at all times. It is suitable for
|
||||
ingesting a large amount of data, essential for the billion scale benchmark.
|
||||
|
||||
Using `memmap_threshold_kb` is not recommended in this case. It would require
|
||||
Using `memmap_threshold` is not recommended in this case. It would require
|
||||
the [optimizer](/documentation/concepts/optimizer/) to constantly
|
||||
transform in-memory segments into memmap segments on disk. This process is
|
||||
slower, and the optimizer can be a bottleneck when ingesting a large amount of
|
||||
+5
-3
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Create and restore from snapshot
|
||||
weight: 14
|
||||
title: Create & Restore Snapshots
|
||||
aliases:
|
||||
- /documentation/tutorials/create-snapshot/
|
||||
weight: 2
|
||||
---
|
||||
|
||||
# Create and restore collections from snapshot
|
||||
# Backup and Restore Qdrant Collections Using Snapshots
|
||||
|
||||
| Time: 20 min | Level: Beginner | | |
|
||||
|--------------|-----------------|--|----|
|
||||
+5
-3
@@ -1,9 +1,11 @@
|
||||
---
|
||||
title: Load Hugging Face dataset
|
||||
weight: 19
|
||||
title: Load a HuggingFace Dataset
|
||||
aliases:
|
||||
- /documentation/tutorials/huggingface-datasets/
|
||||
weight: 3
|
||||
---
|
||||
|
||||
# Loading a dataset from Hugging Face hub
|
||||
# Load and Search Hugging Face Datasets with Qdrant
|
||||
|
||||
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
|
||||
datasets. [Qdrant](https://huggingface.co/Qdrant) also publishes datasets along with the
|
||||
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: Practice Datasets
|
||||
weight: 29
|
||||
partition: build
|
||||
---
|
||||
|
||||
# Common Datasets in Snapshot Format
|
||||
|
||||
+2
-1
@@ -7,4 +7,5 @@ sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
---
|
||||
partition: cloud
|
||||
---
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||
title: "Interfaces & Tools"
|
||||
type: delimiter
|
||||
weight: 25 # Change this weight to change order of sections
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: cloud
|
||||
---
|
||||
+2
-1
@@ -2,9 +2,10 @@
|
||||
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||
title: "Support"
|
||||
type: delimiter
|
||||
weight: 27 # Change this weight to change order of sections
|
||||
weight: 30 # Change this weight to change order of sections
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: cloud
|
||||
---
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||
title: "Essentials"
|
||||
type: delimiter
|
||||
weight: 10 # Change this weight to change order of sections
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: build
|
||||
---
|
||||
+1
@@ -3,6 +3,7 @@
|
||||
title: "Examples"
|
||||
type: delimiter
|
||||
weight: 24 # Change this weight to change order of sections
|
||||
partition: build
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||
title: "Getting Started"
|
||||
type: delimiter
|
||||
weight: 1 # Change this weight to change order of sections
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: qdrant
|
||||
---
|
||||
+1
@@ -7,4 +7,5 @@ sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: build
|
||||
---
|
||||
+1
@@ -7,4 +7,5 @@ sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: cloud
|
||||
---
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||
title: "Support"
|
||||
type: delimiter
|
||||
weight: 30 # Change this weight to change order of sections
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: qdrant
|
||||
---
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||
title: "Tutorials"
|
||||
type: delimiter
|
||||
weight: 15 # Change this weight to change order of sections
|
||||
sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: qdrant
|
||||
---
|
||||
+1
@@ -7,4 +7,5 @@ sitemapExclude: True
|
||||
_build:
|
||||
publishResources: false
|
||||
render: never
|
||||
partition: qdrant
|
||||
---
|
||||
@@ -2,6 +2,7 @@
|
||||
---
|
||||
title: Embeddings
|
||||
weight: 19
|
||||
partition: build
|
||||
---
|
||||
# Supported Embedding Providers & Models
|
||||
|
||||
|
||||
@@ -2,8 +2,8 @@
|
||||
title: Jina Embeddings
|
||||
weight: 1900
|
||||
aliases:
|
||||
- /documentation/embeddings/jina-emebddngs/
|
||||
- ../integrations/jina-embeddings/
|
||||
- /documentation/embeddings/jina-embeddings/
|
||||
- /documentation/integrations/jina-embeddings/
|
||||
---
|
||||
|
||||
# Jina Embeddings
|
||||
@@ -16,13 +16,14 @@ Qdrant users can receive a 10% discount on Jina AI APIs by using the code **QDRA
|
||||
|
||||
| Model | Dimension | Language | MRL (matryoshka) | Context |
|
||||
|:----------------------:|:---------:|:---------:|:-----------:|:---------:|
|
||||
| **jina-clip-v2** | **1024** | **Multilingual (100+, focus on 30)** | **Yes** | **Text/Image** |
|
||||
| jina-embeddings-v3 | 1024 | Multilingual (89 languages) | Yes | 8192 |
|
||||
| jina-embeddings-v2-base-en | 768 | English | No | 8192 |
|
||||
| jina-embeddings-v2-base-de | 768 | German & English | No | 8192 |
|
||||
| jina-embeddings-v2-base-es | 768 | Spanish & English | No | 8192 |
|
||||
| jina-embeddings-v2-base-zh | 768 | Chinese & English | No | 8192 |
|
||||
| jina-embeddings-v2-base-en | 768 | English | No | 8192 |
|
||||
| jina-embeddings-v2-base-de | 768 | German & English | No | 8192 |
|
||||
| jina-embeddings-v2-base-es | 768 | Spanish & English | No | 8192 |
|
||||
| jina-embeddings-v2-base-zh | 768 | Chinese & English | No | 8192 |
|
||||
|
||||
> Jina recommends using `jina-embeddings-v3` as it is the latest and most performant embedding model released by Jina AI.
|
||||
> Jina recommends using `jina-embeddings-v3` for text-only tasks and `jina-clip-v2` for multimodal tasks or when enhanced visual retrieval is required.
|
||||
|
||||
On top of the backbone, `jina-embeddings-v3` has been trained with 5 task-specific adapters for different embedding uses. Include `task` in your request to optimize your downstream application:
|
||||
|
||||
@@ -32,22 +33,22 @@ On top of the backbone, `jina-embeddings-v3` has been trained with 5 task-specif
|
||||
+ **text-matching**: Used to encode text for similarity matching, such as measuring similarity between two sentences.
|
||||
+ **separation**: Used for clustering or reranking tasks.
|
||||
|
||||
`jina-embeddings-v3` supports **Matryoshka Representation Learning**, allowing users to control the embedding dimension with minimal performance loss.
|
||||
`jina-embeddings-v3` and `jina-clip-v2` support **Matryoshka Representation Learning**, allowing users to control the embedding dimension with minimal performance loss.
|
||||
Include `dimensions` in your request to select the desired dimension.
|
||||
By default, **dimensions** is set to 1024, and a number between 256 and 1024 is recommended.
|
||||
You can reference the table below for hints on dimension vs. performance:
|
||||
|
||||
|
||||
| Dimension | 32 | 64 | 128 | 256 | 512 | 768 | 1024 |
|
||||
| Dimension | 32 | 64 | 128 | 256 | 512 | 768 | 1024 |
|
||||
|:----------------------:|:---------:|:---------:|:-----------:|:---------:|:----------:|:---------:|:---------:|
|
||||
| Average Retrieval Performance (nDCG@10) | 52.54 | 58.54 | 61.64 | 62.72 | 63.16 | 63.3 | 63.35 |
|
||||
| Average Retrieval Performance (nDCG@10) | 52.54 | 58.54 | 61.64 | 62.72 | 63.16 | 63.3 | 63.35 |
|
||||
|
||||
`jina-embeddings-v3` supports [Late Chunking](https://jina.ai/news/late-chunking-in-long-context-embedding-models/), the technique to leverage the model's long-context capabilities for generating contextual chunk embeddings. Include `late_chunking=True` in your request to enable contextual chunked representation. When set to true, Jina AI API will concatenate all sentences in the input field and feed them as a single string to the model. Internally, the model embeds this long concatenated string and then performs late chunking, returning a list of embeddings that matches the size of the input list.
|
||||
`jina-embeddings-v3` supports [Late Chunking](https://jina.ai/news/late-chunking-in-long-context-embedding-models/), the technique to leverage the model's long-context capabilities for generating contextual chunk embeddings. Include `late_chunking=True` in your request to enable contextual chunked representation. When set to true, Jina AI API will concatenate all sentences in the input field and feed them as a single string to the model. Internally, the model embeds this long concatenated string and then performs late chunking, returning a list of embeddings that matches the size of the input list.
|
||||
|
||||
## Example
|
||||
|
||||
The code below demonstrate how to use `jina-embeddings-v3` together with Qdrant:
|
||||
### Jina Embeddings v3
|
||||
|
||||
The code below demonstrates how to use `jina-embeddings-v3` with Qdrant:
|
||||
|
||||
```python
|
||||
import requests
|
||||
@@ -59,7 +60,7 @@ from qdrant_client.models import Distance, VectorParams, Batch
|
||||
JINA_API_KEY = "jina_xxxxxxxxxxx"
|
||||
MODEL = "jina-embeddings-v3"
|
||||
DIMENSIONS = 1024 # Or choose your desired output vector dimensionality.
|
||||
TASK = 'retrieval.passage' # For indexing, or set to retrieval.query for quering
|
||||
TASK = 'retrieval.passage' # For indexing, or set to retrieval.query for querying
|
||||
|
||||
# Get embeddings from the API
|
||||
url = "https://api.jina.ai/v1/embeddings"
|
||||
@@ -98,3 +99,99 @@ qdrant_client.upsert(
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
### Jina CLIP v2
|
||||
|
||||
The code below demonstrates how to use `jina-clip-v2` with Qdrant:
|
||||
|
||||
```python
|
||||
import requests
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.models import Distance, VectorParams, PointStruct
|
||||
|
||||
# Provide your Jina API key and choose the model.
|
||||
JINA_API_KEY = "jina_xxxxxxxxxxx"
|
||||
MODEL = "jina-clip-v2"
|
||||
DIMENSIONS = 1024 # Set the desired output vector dimensionality.
|
||||
|
||||
# Define the inputs
|
||||
text_input = "A blue cat"
|
||||
image_url = "https://i.pinimg.com/600x315/21/48/7e/21487e8e0970dd366dafaed6ab25d8d8.jpg"
|
||||
|
||||
# Get embeddings from the Jina API
|
||||
url = "https://api.jina.ai/v1/embeddings"
|
||||
headers = {
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {JINA_API_KEY}",
|
||||
}
|
||||
data = {
|
||||
"input": [
|
||||
{"text": text_input},
|
||||
{"image": image_url},
|
||||
],
|
||||
"model": MODEL,
|
||||
"dimensions": DIMENSIONS,
|
||||
}
|
||||
|
||||
response = requests.post(url, headers=headers, json=data)
|
||||
response_data = response.json()["data"]
|
||||
|
||||
# The model doesn't differentiate between images and text, so we extract output based on the input order.
|
||||
text_embedding = response_data[0]["embedding"]
|
||||
image_embedding = response_data[1]["embedding"]
|
||||
|
||||
# Initialize Qdrant client
|
||||
client = QdrantClient(url="http://localhost:6333/")
|
||||
|
||||
# Create a collection with named vectors
|
||||
collection_name = "MyCollection"
|
||||
client.recreate_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config={
|
||||
"text_vector": VectorParams(size=DIMENSIONS, distance=Distance.DOT),
|
||||
"image_vector": VectorParams(size=DIMENSIONS, distance=Distance.DOT),
|
||||
},
|
||||
)
|
||||
|
||||
client.upsert(
|
||||
collection_name=collection_name,
|
||||
points=[
|
||||
PointStruct(
|
||||
id=0,
|
||||
vector={
|
||||
"text_vector": text_embedding,
|
||||
"image_vector": image_embedding,
|
||||
}
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
# Now let's query the collection
|
||||
search_query = "A purple cat"
|
||||
|
||||
# Get the embedding for the search query from the Jina API
|
||||
url = "https://api.jina.ai/v1/embeddings"
|
||||
headers = {
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {JINA_API_KEY}",
|
||||
}
|
||||
data = {
|
||||
"input": [{"text": search_query}],
|
||||
"model": MODEL,
|
||||
"dimensions": DIMENSIONS,
|
||||
# "task": "retrieval.query" # Uncomment this line for text-to-text retrieval tasks
|
||||
}
|
||||
|
||||
response = requests.post(url, headers=headers, json=data)
|
||||
query_embedding = response.json()["data"][0]["embedding"]
|
||||
|
||||
search_results = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
using="image_vector",
|
||||
limit=5
|
||||
).points
|
||||
|
||||
for result in search_results:
|
||||
print(f"ID: {result.id}, Score: {result.score}")
|
||||
```
|
||||
|
||||
@@ -3,45 +3,54 @@ title: Ollama
|
||||
weight: 2600
|
||||
---
|
||||
|
||||
# Using Ollama with Qdrant
|
||||
|
||||
Ollama provides specialized embeddings for niche applications. Ollama supports a variety of embedding models, making it possible to build retrieval augmented generation (RAG) applications that combine text prompts with existing documents or other data in specialized areas.
|
||||
|
||||
# Using Ollama with Qdrant
|
||||
|
||||
[Ollama](https://ollama.com) provides specialized embeddings for niche applications. Ollama supports a [variety of embedding models](https://ollama.com/search?c=embedding), making it possible to build retrieval augmented generation (RAG) applications that combine text prompts with existing documents or other data in specialized areas.
|
||||
|
||||
## Installation
|
||||
|
||||
You can install the required package using the following pip command:
|
||||
You can install the required packages using the following pip command:
|
||||
|
||||
```bash
|
||||
pip install ollama
|
||||
pip install ollama qdrant-client
|
||||
```
|
||||
|
||||
## Integration Example
|
||||
|
||||
The following code assumes Ollama is accessible at port `11434` and Qdrant at port `6334`.
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
from qdrant_client.models import Batch
|
||||
from ollama import Ollama
|
||||
from qdrant_client import QdrantClient, models
|
||||
import ollama
|
||||
|
||||
# Initialize Ollama model
|
||||
model = Ollama("ollama-unique")
|
||||
COLLECTION_NAME = "NicheApplications"
|
||||
|
||||
# Generate embeddings for niche applications
|
||||
text = "Ollama excels in niche applications with specific embeddings."
|
||||
embeddings = model.embed(text)
|
||||
# Initialize Ollama client
|
||||
oclient = ollama.Client(host="localhost")
|
||||
|
||||
# Initialize Qdrant client
|
||||
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
|
||||
qclient = QdrantClient(host="localhost", port=6333)
|
||||
|
||||
# Upsert the embedding into Qdrant
|
||||
qdrant_client.upsert(
|
||||
collection_name="NicheApplications",
|
||||
points=Batch(
|
||||
ids=[1],
|
||||
vectors=[embeddings],
|
||||
# Text to embed
|
||||
text = "Ollama excels in niche applications with specific embeddings"
|
||||
|
||||
# Generate embeddings
|
||||
response = oclient.embeddings(model="llama3.2", prompt=text)
|
||||
embeddings = response["embedding"]
|
||||
|
||||
# Create a collection if it doesn't already exist
|
||||
if not qclient.collection_exists(COLLECTION_NAME):
|
||||
qclient.create_collection(
|
||||
collection_name=COLLECTION_NAME,
|
||||
vectors_config=models.VectorParams(
|
||||
size=len(embeddings), distance=models.Distance.COSINE
|
||||
),
|
||||
)
|
||||
|
||||
# Upload the vectors to the collection along with the original text as payload
|
||||
qclient.upsert(
|
||||
collection_name=COLLECTION_NAME,
|
||||
points=[models.PointStruct(id=1, vector=embeddings, payload={"text": text})],
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: Build Prototypes
|
||||
weight: 26
|
||||
partition: build
|
||||
---
|
||||
# Examples
|
||||
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
---
|
||||
title: FAQ
|
||||
weight: 28
|
||||
weight: 31
|
||||
partition: qdrant
|
||||
is_empty: true
|
||||
---
|
||||
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: "FastEmbed"
|
||||
weight: 6
|
||||
weight: 7
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# What is FastEmbed?
|
||||
|
||||
@@ -5,25 +5,54 @@ weight: 6
|
||||
|
||||
# How to Generate ColBERT Multivectors with FastEmbed
|
||||
|
||||
With FastEmbed, you can use ColBERT to generate multivector embeddings. ColBERT is a powerful retrieval model that combines the strength of BERT embeddings with efficient late interaction techniques. FastEmbed will provide you with an optimized pipeline to utilize these embeddings in your search tasks.
|
||||
## ColBERT
|
||||
|
||||
Please note that ColBERT requires more resources than other no-interaction models. We recommend you use ColBERT as a re-ranker instead of a first-stage retriever.
|
||||
ColBERT is an embedding model that produces a matrix (multivector) representation of input text,
|
||||
generating one vector per token (a token being a meaningful text unit for a machine learning model).
|
||||
This approach allows ColBERT to capture more nuanced input semantics than many dense embedding models,
|
||||
which represent an entire input with a single vector. By producing more granular input representations,
|
||||
ColBERT becomes a strong retriever. However, this advantage comes at the cost of increased resource consumption compared to
|
||||
traditional dense embedding models, both in terms of speed and memory.
|
||||
|
||||
The first-stage retriever can retrieve 100-500 examples. This task would be done by a simpler model. Then, you can rank the leftover results with ColBERT.
|
||||
Despite ColBERT being a powerful retriever, it's speed limitation might make it less suitable for large-scale retrieval.
|
||||
Therefore, we generally recommend using ColBERT for reranking a small set of already retrieved examples, rather than for first-stage retrieval.
|
||||
A simple dense retriever can initially retrieve around 100-500 candidates, which can then be reranked with ColBERT to bring the most relevant results
|
||||
to the top.
|
||||
|
||||
ColBERT is a considerable alternative of a reranking model to [cross-encoders](https://sbert.net/examples/applications/cross-encoder/README.html), since
|
||||
It tends to be faster on inference time due to its `late interaction` mechanism.
|
||||
|
||||
How does `late interaction` work? Cross-encoders ingest a query and a document glued together as one input.
|
||||
A cross-encoder model divides this input into meaningful (for the model) parts and checks how these parts relate.
|
||||
So, all interactions between the query and the document happen "early" inside the model.
|
||||
Late interaction models, such as ColBERT, only do the first part, generating document and query parts suitable for comparison.
|
||||
All interactions between these parts are expected to be done "later" outside the model.
|
||||
|
||||
## Using ColBERT in Qdrant
|
||||
|
||||
Qdrant supports [multivector representations](https://qdrant.tech/documentation/concepts/vectors/#multivectors) out of the box so that you can use any late interaction model as `ColBERT` or `ColPali` in Qdrant without any additional pre/post-processing.
|
||||
|
||||
This tutorial uses ColBERT as a first-stage retriever on a toy dataset.
|
||||
You can see how to use ColBERT as a reranker in our [multi-stage queries documentation](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries).
|
||||
## Setup
|
||||
|
||||
This command imports all late interaction models for text embedding.
|
||||
Install `fastembed`.
|
||||
|
||||
```python
|
||||
pip install fastembed
|
||||
```
|
||||
|
||||
Imports late interaction models for text embedding.
|
||||
|
||||
```python
|
||||
from fastembed import LateInteractionTextEmbedding
|
||||
```
|
||||
You can list which models are supported in your version of FastEmbed.
|
||||
You can list which late interaction models are supported in FastEmbed.
|
||||
|
||||
```python
|
||||
LateInteractionTextEmbedding.list_supported_models()
|
||||
```
|
||||
This command displays the available models. The output shows details about the ColBERT model, including its dimensions, description, size, sources, and model file.
|
||||
This command displays the available models. The output shows details about the model, including output embedding dimensions, model description, model size, model sources, and model file.
|
||||
|
||||
```python
|
||||
[{'model': 'colbert-ir/colbertv2.0',
|
||||
@@ -31,7 +60,13 @@ This command displays the available models. The output shows details about the C
|
||||
'description': 'Late interaction model',
|
||||
'size_in_GB': 0.44,
|
||||
'sources': {'hf': 'colbert-ir/colbertv2.0'},
|
||||
'model_file': 'model.onnx'}]
|
||||
'model_file': 'model.onnx'},
|
||||
{'model': 'answerdotai/answerai-colbert-small-v1',
|
||||
'dim': 96,
|
||||
'description': 'Text embeddings, Unimodal (text), Multilingual (~100 languages), 512 input tokens truncation, 2024 year',
|
||||
'size_in_GB': 0.13,
|
||||
'sources': {'hf': 'answerdotai/answerai-colbert-small-v1'},
|
||||
'model_file': 'vespa_colbert.onnx'}]
|
||||
```
|
||||
Now, load the model.
|
||||
|
||||
@@ -42,105 +77,153 @@ The model files will be fetched and downloaded, with progress showing.
|
||||
|
||||
## Embed data
|
||||
|
||||
First, you need to define both documents and queries.
|
||||
We will vectorize a toy movie description dataset with ColBERT:
|
||||
|
||||
<details>
|
||||
<summary> <span style="background-color: gray; color: black;"> Movie description dataset </span> </summary>
|
||||
|
||||
```python
|
||||
documents = [
|
||||
"ColBERT is a late interaction text embedding model, however, there are also other models such as TwinBERT.",
|
||||
"On the contrary to the late interaction models, the early interaction models contains interaction steps at embedding generation process",
|
||||
]
|
||||
queries = [
|
||||
"Are there any other late interaction text embedding models except ColBERT?",
|
||||
"What is the difference between late interaction and early interaction text embedding models?",
|
||||
]
|
||||
descriptions = ["In 1431, Jeanne d'Arc is placed on trial on charges of heresy. The ecclesiastical jurists attempt to force Jeanne to recant her claims of holy visions.",
|
||||
"A film projectionist longs to be a detective, and puts his meagre skills to work when he is framed by a rival for stealing his girlfriend's father's pocketwatch.",
|
||||
"A group of high-end professional thieves start to feel the heat from the LAPD when they unknowingly leave a clue at their latest heist.",
|
||||
"A petty thief with an utter resemblance to a samurai warlord is hired as the lord's double. When the warlord later dies the thief is forced to take up arms in his place.",
|
||||
"A young boy named Kubo must locate a magical suit of armour worn by his late father in order to defeat a vengeful spirit from the past.",
|
||||
"A biopic detailing the 2 decades that Punjabi Sikh revolutionary Udham Singh spent planning the assassination of the man responsible for the Jallianwala Bagh massacre.",
|
||||
"When a machine that allows therapists to enter their patients' dreams is stolen, all hell breaks loose. Only a young female therapist, Paprika, can stop it.",
|
||||
"An ordinary word processor has the worst night of his life after he agrees to visit a girl in Soho whom he met that evening at a coffee shop.",
|
||||
"A story that revolves around drug abuse in the affluent north Indian State of Punjab and how the youth there have succumbed to it en-masse resulting in a socio-economic decline.",
|
||||
"A world-weary political journalist picks up the story of a woman's search for her son, who was taken away from her decades ago after she became pregnant and was forced to live in a convent.",
|
||||
"Concurrent theatrical ending of the TV series Neon Genesis Evangelion (1995).",
|
||||
"During World War II, a rebellious U.S. Army Major is assigned a dozen convicted murderers to train and lead them into a mass assassination mission of German officers.",
|
||||
"The toys are mistakenly delivered to a day-care center instead of the attic right before Andy leaves for college, and it's up to Woody to convince the other toys that they weren't abandoned and to return home.",
|
||||
"A soldier fighting aliens gets to relive the same day over and over again, the day restarting every time he dies.",
|
||||
"After two male musicians witness a mob hit, they flee the state in an all-female band disguised as women, but further complications set in.",
|
||||
"Exiled into the dangerous forest by her wicked stepmother, a princess is rescued by seven dwarf miners who make her part of their household.",
|
||||
"A renegade reporter trailing a young runaway heiress for a big story joins her on a bus heading from Florida to New York, and they end up stuck with each other when the bus leaves them behind at one of the stops.",
|
||||
"Story of 40-man Turkish task force who must defend a relay station.",
|
||||
"Spinal Tap, one of England's loudest bands, is chronicled by film director Marty DiBergi on what proves to be a fateful tour.",
|
||||
"Oskar, an overlooked and bullied boy, finds love and revenge through Eli, a beautiful but peculiar girl."]
|
||||
```
|
||||
</details>
|
||||
|
||||
**Note:** ColBERT computes document and query embeddings differently. Make sure to use the corresponding methods.
|
||||
|
||||
Now, create embeddings from both documents and queries.
|
||||
The vectorization is done with an `embed` generator function.
|
||||
|
||||
```python
|
||||
document_embeddings = list(
|
||||
embedding_model.embed(documents)
|
||||
) # embed and qury_embed return generators,
|
||||
# which we need to evaluate by writing them to a list
|
||||
query_embeddings = list(embedding_model.query_embed(queries))
|
||||
|
||||
```
|
||||
Display the shapes of document and query embeddings.
|
||||
|
||||
```python
|
||||
document_embeddings[0].shape, query_embeddings[0].shape
|
||||
```
|
||||
|
||||
You should get something like this:
|
||||
|
||||
```python
|
||||
((26, 128), (32, 128))
|
||||
```
|
||||
|
||||
Don't worry about query embeddings having the bigger shape in this case. ColBERT authors recommend to pad queries with [MASK] tokens to 32 tokens. They also recommend truncating queries to 32 tokens, however, we don't do that in FastEmbed so that you can put some straight into the queries.
|
||||
|
||||
## Compute similarity
|
||||
|
||||
This function calculates the relevance scores using the MaxSim operator, sorts the documents based on these scores, and returns the indices of the top-k documents.
|
||||
|
||||
```python
|
||||
import numpy as np
|
||||
|
||||
|
||||
def compute_relevance_scores(query_embedding: np.array, document_embeddings: np.array, k: int):
|
||||
"""
|
||||
Compute relevance scores for top-k documents given a query.
|
||||
|
||||
:param query_embedding: Numpy array representing the query embedding, shape: [num_query_terms, embedding_dim]
|
||||
:param document_embeddings: Numpy array representing embeddings for documents, shape: [num_documents, max_doc_length, embedding_dim]
|
||||
:param k: Number of top documents to return
|
||||
:return: Indices of the top-k documents based on their relevance scores
|
||||
"""
|
||||
# Compute batch dot-product of query_embedding and document_embeddings
|
||||
# Resulting shape: [num_documents, num_query_terms, max_doc_length]
|
||||
scores = np.matmul(query_embedding, document_embeddings.transpose(0, 2, 1))
|
||||
|
||||
# Apply max-pooling across document terms (axis=2) to find the max similarity per query term
|
||||
# Shape after max-pool: [num_documents, num_query_terms]
|
||||
max_scores_per_query_term = np.max(scores, axis=2)
|
||||
|
||||
# Sum the scores across query terms to get the total score for each document
|
||||
# Shape after sum: [num_documents]
|
||||
total_scores = np.sum(max_scores_per_query_term, axis=1)
|
||||
|
||||
# Sort the documents based on their total scores and get the indices of the top-k documents
|
||||
sorted_indices = np.argsort(total_scores)[::-1][:k]
|
||||
|
||||
return sorted_indices
|
||||
```
|
||||
Calculate sorted indices.
|
||||
|
||||
```python
|
||||
sorted_indices = compute_relevance_scores(
|
||||
np.array(query_embeddings[0]), np.array(document_embeddings), k=3
|
||||
descriptions_embeddings = list(
|
||||
embedding_model.embed(descriptions)
|
||||
)
|
||||
print("Sorted document indices:", sorted_indices)
|
||||
```
|
||||
The output shows the sorted document indices based on the relevance to the query.
|
||||
Let's check the size of one of the produced embeddings.
|
||||
|
||||
```python
|
||||
Sorted document indices: [0 1]
|
||||
descriptions_embeddings[0].shape
|
||||
```
|
||||
|
||||
## Show results
|
||||
|
||||
```python
|
||||
print(f"Query: {queries[0]}")
|
||||
for index in sorted_indices:
|
||||
print(f"Document: {documents[index]}")
|
||||
```
|
||||
|
||||
The query and corresponding sorted documents are displayed, showing the relevance of each document to the query.
|
||||
We get the following result
|
||||
|
||||
```bash
|
||||
Query: Are there any other late interaction text embedding models except ColBERT?
|
||||
Document: ColBERT is a late interaction text embedding model, however, there are also other models such as TwinBERT.
|
||||
Document: On the contrary to the late interaction models, the early interaction models contains interaction steps at embedding generation process
|
||||
(48, 128)
|
||||
```
|
||||
That means that for the first description, we have **48** vectors of lengths **128** representing it.
|
||||
|
||||
## Upload embeddings to Qdrant
|
||||
|
||||
Install `qdrant-client`
|
||||
|
||||
```python
|
||||
pip install qdrant-client
|
||||
```
|
||||
|
||||
Qdrant Client has a simple in-memory mode that allows you to experiment locally on small data volumes.
|
||||
Alternatively, you could use for experiments [a free cluster](https://qdrant.tech/documentation/cloud/create-cluster/#create-a-cluster) in Qdrant Cloud.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
qdrant_client = QdrantClient(":memory:") # Qdrant is running from RAM.
|
||||
```
|
||||
|
||||
Now, let's create a small [collection](https://qdrant.tech/documentation/concepts/collections/) with our movie data.
|
||||
For that, we will use the [multivectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors) functionality supported in Qdrant.
|
||||
To configure multivector collection, we need to specify:
|
||||
- similarity metric between vectors;
|
||||
- the size of each vector (for ColBERT, it's **128**);
|
||||
- similarity metric between multivectors (matrices), for example, `maximum`, so for vector from matrix A, we find the most similar vector from matrix B, and their similarity score will be out matrix similarity.
|
||||
|
||||
```python
|
||||
qdrant_client.create_collection(
|
||||
collection_name="movies",
|
||||
vectors_config=models.VectorParams(
|
||||
size=128, #size of each vector produced by ColBERT
|
||||
distance=models.Distance.COSINE, #similarity metric between each vector
|
||||
multivector_config=models.MultiVectorConfig(
|
||||
comparator=models.MultiVectorComparator.MAX_SIM #similarity metric between multivectors (matrices)
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
To make this collection human-readable, let's save movie metadata (name, description in text form and movie's length) together with an embedded description.
|
||||
|
||||
<details>
|
||||
<summary> <span style="background-color: gray; color: black;"> Movie metadata </span> </summary>
|
||||
|
||||
```python
|
||||
metadata = [{"movie_name": "The Passion of Joan of Arc", "movie_watch_time_min": 114, "movie_description": "In 1431, Jeanne d'Arc is placed on trial on charges of heresy. The ecclesiastical jurists attempt to force Jeanne to recant her claims of holy visions."},
|
||||
{"movie_name": "Sherlock Jr.", "movie_watch_time_min": 45, "movie_description": "A film projectionist longs to be a detective, and puts his meagre skills to work when he is framed by a rival for stealing his girlfriend's father's pocketwatch."},
|
||||
{"movie_name": "Heat", "movie_watch_time_min": 170, "movie_description": "A group of high-end professional thieves start to feel the heat from the LAPD when they unknowingly leave a clue at their latest heist."},
|
||||
{"movie_name": "Kagemusha", "movie_watch_time_min": 162, "movie_description": "A petty thief with an utter resemblance to a samurai warlord is hired as the lord's double. When the warlord later dies the thief is forced to take up arms in his place."},
|
||||
{"movie_name": "Kubo and the Two Strings", "movie_watch_time_min": 101, "movie_description": "A young boy named Kubo must locate a magical suit of armour worn by his late father in order to defeat a vengeful spirit from the past."},
|
||||
{"movie_name": "Sardar Udham", "movie_watch_time_min": 164, "movie_description": "A biopic detailing the 2 decades that Punjabi Sikh revolutionary Udham Singh spent planning the assassination of the man responsible for the Jallianwala Bagh massacre."},
|
||||
{"movie_name": "Paprika", "movie_watch_time_min": 90, "movie_description": "When a machine that allows therapists to enter their patients' dreams is stolen, all hell breaks loose. Only a young female therapist, Paprika, can stop it."},
|
||||
{"movie_name": "After Hours", "movie_watch_time_min": 97, "movie_description": "An ordinary word processor has the worst night of his life after he agrees to visit a girl in Soho whom he met that evening at a coffee shop."},
|
||||
{"movie_name": "Udta Punjab", "movie_watch_time_min": 148, "movie_description": "A story that revolves around drug abuse in the affluent north Indian State of Punjab and how the youth there have succumbed to it en-masse resulting in a socio-economic decline."},
|
||||
{"movie_name": "Philomena", "movie_watch_time_min": 98, "movie_description": "A world-weary political journalist picks up the story of a woman's search for her son, who was taken away from her decades ago after she became pregnant and was forced to live in a convent."},
|
||||
{"movie_name": "Neon Genesis Evangelion: The End of Evangelion", "movie_watch_time_min": 87, "movie_description": "Concurrent theatrical ending of the TV series Neon Genesis Evangelion (1995)."},
|
||||
{"movie_name": "The Dirty Dozen", "movie_watch_time_min": 150, "movie_description": "During World War II, a rebellious U.S. Army Major is assigned a dozen convicted murderers to train and lead them into a mass assassination mission of German officers."},
|
||||
{"movie_name": "Toy Story 3", "movie_watch_time_min": 103, "movie_description": "The toys are mistakenly delivered to a day-care center instead of the attic right before Andy leaves for college, and it's up to Woody to convince the other toys that they weren't abandoned and to return home."},
|
||||
{"movie_name": "Edge of Tomorrow", "movie_watch_time_min": 113, "movie_description": "A soldier fighting aliens gets to relive the same day over and over again, the day restarting every time he dies."},
|
||||
{"movie_name": "Some Like It Hot", "movie_watch_time_min": 121, "movie_description": "After two male musicians witness a mob hit, they flee the state in an all-female band disguised as women, but further complications set in."},
|
||||
{"movie_name": "Snow White and the Seven Dwarfs", "movie_watch_time_min": 83, "movie_description": "Exiled into the dangerous forest by her wicked stepmother, a princess is rescued by seven dwarf miners who make her part of their household."},
|
||||
{"movie_name": "It Happened One Night", "movie_watch_time_min": 105, "movie_description": "A renegade reporter trailing a young runaway heiress for a big story joins her on a bus heading from Florida to New York, and they end up stuck with each other when the bus leaves them behind at one of the stops."},
|
||||
{"movie_name": "Nefes: Vatan Sagolsun", "movie_watch_time_min": 128, "movie_description": "Story of 40-man Turkish task force who must defend a relay station."},
|
||||
{"movie_name": "This Is Spinal Tap", "movie_watch_time_min": 82, "movie_description": "Spinal Tap, one of England's loudest bands, is chronicled by film director Marty DiBergi on what proves to be a fateful tour."},
|
||||
{"movie_name": "Let the Right One In", "movie_watch_time_min": 114, "movie_description": "Oskar, an overlooked and bullied boy, finds love and revenge through Eli, a beautiful but peculiar girl."}]
|
||||
```
|
||||
</details>
|
||||
|
||||
```python
|
||||
qdrant_client.upload_points(
|
||||
collection_name="movies",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=idx,
|
||||
payload=metadata[idx],
|
||||
vector=vector
|
||||
)
|
||||
for idx, vector in enumerate(descriptions_embeddings)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
## Querying
|
||||
|
||||
ColBERT uses two distinct methods for embedding documents and queries, as do we in Fastembed. However, we altered query pre-processing used in ColBERT, so we don't have to cut all queries after 32-token length but ingest longer queries directly.
|
||||
|
||||
```python
|
||||
qdrant_client.query_points(
|
||||
collection_name="movies",
|
||||
query=list(embedding_model.query_embed("A movie for kids with fantasy elements and wonders"))[0], #converting generator object into numpy.ndarray
|
||||
limit=1, #How many closest to the query movies we would like to get
|
||||
#with_vectors=True, #If this option is used, vectors will also be returned
|
||||
with_payload=True #So metadata is provided in the output
|
||||
)
|
||||
```
|
||||
|
||||
The result is the following:
|
||||
|
||||
```bash
|
||||
QueryResponse(points=[ScoredPoint(id=4, version=0, score=12.063469,
|
||||
payload={'movie_name': 'Kubo and the Two Strings', 'movie_watch_time_min': 101,
|
||||
'movie_description': 'A young boy named Kubo must locate a magical suit of armour worn by his late father in order to defeat a vengeful spirit from the past.'},
|
||||
vector=None, shard_key=None, order_value=None)])
|
||||
```
|
||||
|
||||
@@ -1,28 +1,39 @@
|
||||
---
|
||||
title: Frameworks
|
||||
weight: 20
|
||||
partition: build
|
||||
---
|
||||
|
||||
## Framework Integrations
|
||||
|
||||
| Framework | Description |
|
||||
| ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
||||
| [AutoGen](/documentation/frameworks/autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
|
||||
| [Canopy](/documentation/frameworks/canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
|
||||
| [Cheshire Cat](/documentation/frameworks/cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
|
||||
| [DocArray](/documentation/frameworks/docarray/) | Python library for managing data in multi-modal AI applications. |
|
||||
| [DSPy](/documentation/frameworks/dspy/) | Framework for algorithmically optimizing LM prompts and weights. |
|
||||
| [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. |
|
||||
| [Genkit](/documentation/frameworks/genkit/) | Framework to build, deploy, and monitor production-ready AI-powered apps. |
|
||||
| [Haystack](/documentation/frameworks/haystack/) | LLM orchestration framework to build customizable, production-ready LLM applications. |
|
||||
| [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [Langchain4j](/documentation/frameworks/langchain4j/) | Java framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
|
||||
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
|
||||
| [MemGPT](/documentation/frameworks/memgpt/) | System to build LLM agents with long term memory & custom tools |
|
||||
| [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language |
|
||||
| [Semantic Router](/documentation/frameworks/semantic-router/) | Python library to build a decision-making layer for AI applications using vector search. |
|
||||
| [Spring AI](/documentation/frameworks/spring-ai/) | Java AI framework for building with Spring design principles such as portability and modular design. |
|
||||
| [txtai](/documentation/frameworks/txtai/) | Python library for semantic search, LLM orchestration and language model workflows. |
|
||||
| [Vanna AI](/documentation/frameworks/vanna-ai/) | Python RAG framework for SQL generation and querying. |
|
||||
| Framework | Description |
|
||||
| ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
|
||||
| [AutoGen](/documentation/frameworks/autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
|
||||
| [Canopy](/documentation/frameworks/canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
|
||||
| [Cheshire Cat](/documentation/frameworks/cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
|
||||
| [CrewAI](/documentation/frameworks/crewai/) | CrewAI is a framework to build automated workflows using multiple AI agents that perform complex tasks. |
|
||||
| [DocArray](/documentation/frameworks/docarray/) | Python library for managing data in multi-modal AI applications. |
|
||||
| [DSPy](/documentation/frameworks/dspy/) | Framework for algorithmically optimizing LM prompts and weights. |
|
||||
| [Feast](/documentation/frameworks/feast/) | Open-source feature store to operate production ML systems at scale as a set of features. |
|
||||
| [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. |
|
||||
| [Genkit](/documentation/frameworks/genkit/) | Framework to build, deploy, and monitor production-ready AI-powered apps. |
|
||||
| [Haystack](/documentation/frameworks/haystack/) | LLM orchestration framework to build customizable, production-ready LLM applications. |
|
||||
| [Lakechain](/documentation/frameworks/lakechain/) | Python framework for deploying document processing pipelines on AWS using infrastructure-as-code. |
|
||||
| [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [Langchain4j](/documentation/frameworks/langchain4j/) | Java framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [LangGraph](/documentation/frameworks/langgraph/) | Python, Javascript libraries for building stateful, multi-actor applications. |
|
||||
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
|
||||
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
|
||||
| [MemGPT](/documentation/frameworks/memgpt/) | System to build LLM agents with long term memory & custom tools |
|
||||
| [Neo4j GraphRAG](/documentation/frameworks/neo4j-graphrag/) | Package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. |
|
||||
| [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language |
|
||||
| [Ragbits](/documentation/frameworks/ragbits/) | Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. |
|
||||
| [Rig-rs](/documentation/frameworks/rig-rs/) | Rust library for building scalable, modular, and ergonomic LLM-powered applications. |
|
||||
| [Semantic Router](/documentation/frameworks/semantic-router/) | Python library to build a decision-making layer for AI applications using vector search. |
|
||||
| [Spring AI](/documentation/frameworks/spring-ai/) | Java AI framework for building with Spring design principles such as portability and modular design. |
|
||||
| [Swarm](/documentation/frameworks/swarm/) | Python framework for managing multiple AI agents that can work together. |
|
||||
| [Sycamore](/documentation/frameworks/sycamore/) | Document processing engine for ETL, RAG, LLM-based applications, and analytics on unstructured data. |
|
||||
| [Testcontainers](/documentation/frameworks/testcontainers/) | Framework for providing throwaway, lightweight instances of systems for testing |
|
||||
| [txtai](/documentation/frameworks/txtai/) | Python library for semantic search, LLM orchestration and language model workflows. |
|
||||
| [Vanna AI](/documentation/frameworks/vanna-ai/) | Python RAG framework for SQL generation and querying. |
|
||||
|
||||
@@ -5,100 +5,100 @@ aliases: [ ../integrations/autogen/ ]
|
||||
|
||||
# Microsoft Autogen
|
||||
|
||||
[AutoGen](https://github.com/microsoft/autogen) is a framework that enables the development of LLM applications using multiple agents that can converse with each other to solve tasks. AutoGen agents are customizable, conversable, and seamlessly allow human participation. They can operate in various modes that employ combinations of LLMs, human inputs, and tools.
|
||||
[AutoGen](https://github.com/microsoft/autogen/tree/0.2) is an open-source programming framework for building AI agents and facilitating cooperation among multiple agents to solve tasks.
|
||||
|
||||
- Multi-agent conversations: AutoGen agents can communicate with each other to solve tasks. This allows for more complex and sophisticated applications than would be possible with a single LLM.
|
||||
- Customization: AutoGen agents can be customized to meet the specific needs of an application. This includes the ability to choose the LLMs to use, the types of human input to allow, and the tools to employ.
|
||||
- Human participation: AutoGen seamlessly allows human participation. This means that humans can provide input and feedback to the agents as needed.
|
||||
|
||||
With the Autogen-Qdrant integration, you can use the `QdrantRetrieveUserProxyAgent` from autogen to build retrieval augmented generation(RAG) services with ease.
|
||||
- Customization: AutoGen agents can be customized to meet the specific needs of an application. This includes the ability to choose the LLMs to use, the types of human input to allow, and the tools to employ.
|
||||
|
||||
- Human participation: AutoGen allows human participation. This means that humans can provide input and feedback to the agents as needed.
|
||||
|
||||
With the [Autogen-Qdrant integration](https://microsoft.github.io/autogen/0.2/docs/reference/agentchat/contrib/vectordb/qdrant/), you build Autogen workflows backed by Qdrant't performant retrievals.
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
pip install "pyautogen[retrievechat]" "qdrant_client[fastembed]"
|
||||
pip install "autogen-agentchat[retrievechat-qdrant]"
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
A demo application that generates code based on context w/o human feedback
|
||||
|
||||
#### Set your API Endpoint
|
||||
|
||||
The config_list_from_json function loads a list of configurations from an environment variable or a JSON file.
|
||||
#### Configuration
|
||||
|
||||
```python
|
||||
from autogen import config_list_from_json
|
||||
from autogen.agentchat.contrib.retrieve_assistant_agent import RetrieveAssistantAgent
|
||||
from autogen.agentchat.contrib.qdrant_retrieve_user_proxy_agent import QdrantRetrieveUserProxyAgent
|
||||
from qdrant_client import QdrantClient
|
||||
import autogen
|
||||
|
||||
config_list = config_list_from_json(
|
||||
env_or_file="OAI_CONFIG_LIST",
|
||||
file_location="."
|
||||
)
|
||||
config_list = autogen.config_list_from_json("OAI_CONFIG_LIST")
|
||||
```
|
||||
|
||||
It first looks for the environment variable "OAI_CONFIG_LIST" which needs to be a valid JSON string. If that variable is not found, it then looks for a JSON file named "OAI_CONFIG_LIST". The file structure sample can be found [here](https://github.com/microsoft/autogen/blob/main/OAI_CONFIG_LIST_sample).
|
||||
The `config_list_from_json` function first looks for the environment variable `OAI_CONFIG_LIST` which needs to be a valid JSON string. If not found, it then looks for a JSON file named `OAI_CONFIG_LIST`. A sample file can be found [here](https://github.com/microsoft/autogen/blob/0.2/OAI_CONFIG_LIST_sample).
|
||||
|
||||
#### Construct agents for RetrieveChat
|
||||
|
||||
We start by initializing the RetrieveAssistantAgent and QdrantRetrieveUserProxyAgent. The system message needs to be set to "You are a helpful assistant." for RetrieveAssistantAgent. The detailed instructions are given in the user message.
|
||||
|
||||
```python
|
||||
# Print the generation steps
|
||||
autogen.ChatCompletion.start_logging()
|
||||
from qdrant_client import QdrantClient
|
||||
from sentence_transformers import SentenceTransformer
|
||||
|
||||
# 1. create a RetrieveAssistantAgent instance named "assistant"
|
||||
assistant = RetrieveAssistantAgent(
|
||||
from autogen import AssistantAgent
|
||||
from autogen.agentchat.contrib.retrieve_user_proxy_agent import RetrieveUserProxyAgent
|
||||
|
||||
# 1. Create an AssistantAgent instance named "assistant"
|
||||
assistant = AssistantAgent(
|
||||
name="assistant",
|
||||
system_message="You are a helpful assistant.",
|
||||
llm_config={
|
||||
"request_timeout": 600,
|
||||
"seed": 42,
|
||||
"timeout": 600,
|
||||
"cache_seed": 42,
|
||||
"config_list": config_list,
|
||||
},
|
||||
)
|
||||
|
||||
# 2. create a QdrantRetrieveUserProxyAgent instance named "qdrantagent"
|
||||
# By default, the human_input_mode is "ALWAYS", i.e. the agent will ask for human input at every step.
|
||||
# `docs_path` is the path to the docs directory.
|
||||
# `task` indicates the kind of task we're working on.
|
||||
# `chunk_token_size` is the chunk token size for the retrieve chat.
|
||||
# We use an in-memory QdrantClient instance here. Not recommended for production.
|
||||
sentence_transformer_ef = SentenceTransformer("all-distilroberta-v1").encode
|
||||
client = QdrantClient(url="http://localhost:6333/")
|
||||
|
||||
rag_proxy_agent = QdrantRetrieveUserProxyAgent(
|
||||
name="qdrantagent",
|
||||
# 2. Create the RetrieveUserProxyAgent instance named "ragproxyagent"
|
||||
# Refer to https://microsoft.github.io/autogen/docs/reference/agentchat/contrib/retrieve_user_proxy_agent
|
||||
# for more information on the RetrieveUserProxyAgent
|
||||
ragproxyagent = RetrieveUserProxyAgent(
|
||||
name="ragproxyagent",
|
||||
human_input_mode="NEVER",
|
||||
max_consecutive_auto_reply=10,
|
||||
retrieve_config={
|
||||
"task": "code",
|
||||
"docs_path": "./path/to/docs",
|
||||
"docs_path": [
|
||||
"path/to/some/doc.md",
|
||||
"path/to/some/other/doc.md",
|
||||
],
|
||||
"chunk_token_size": 2000,
|
||||
"model": config_list[0]["model"],
|
||||
"client": QdrantClient(":memory:"),
|
||||
"embedding_model": "BAAI/bge-small-en-v1.5",
|
||||
"vector_db": "qdrant",
|
||||
"db_config": {"client": client},
|
||||
"get_or_create": True,
|
||||
"overwrite": True,
|
||||
"embedding_function": sentence_transformer_ef, # Defaults to "BAAI/bge-small-en-v1.5" via FastEmbed
|
||||
},
|
||||
code_execution_config=False,
|
||||
)
|
||||
```
|
||||
|
||||
#### Run the retriever service
|
||||
#### Run the agent
|
||||
|
||||
```python
|
||||
# Always reset the assistant before starting a new conversation.
|
||||
assistant.reset()
|
||||
|
||||
# We use the ragproxyagent to generate a prompt to be sent to the assistant as the initial message.
|
||||
# The assistant receives the message and generates a response. The response will be sent back to the ragproxyagent for processing.
|
||||
# The conversation continues until the termination condition is met, in RetrieveChat, the termination condition when no human-in-loop is no code block detected.
|
||||
# The assistant receives it and generates a response. The response will be sent back to the ragproxyagent for processing.
|
||||
# The conversation continues until the termination condition is met.
|
||||
|
||||
# The query used below is for demonstration. It should usually be related to the docs made available to the agent
|
||||
code_problem = "How can I use FLAML to perform a classification task?"
|
||||
rag_proxy_agent.initiate_chat(assistant, problem=code_problem)
|
||||
qa_problem = "What is the .....?"
|
||||
chat_results = ragproxyagent.initiate_chat(assistant, message=ragproxyagent.message_generator, problem=qa_problem)
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- Autogen [examples](https://microsoft.github.io/autogen/docs/Examples)
|
||||
- AutoGen [documentation](https://microsoft.github.io/autogen/)
|
||||
- [Source Code](https://github.com/microsoft/autogen/blob/main/autogen/agentchat/contrib/qdrant_retrieve_user_proxy_agent.py)
|
||||
- AutoGen [documentation](https://microsoft.github.io/autogen/0.2)
|
||||
- Autogen [examples](https://microsoft.github.io/autogen/0.2/docs/Examples)
|
||||
- [Source Code](https://github.com/microsoft/autogen/blob/0.2/autogen/agentchat/contrib/vectordb/qdrant.py)
|
||||
|
||||
@@ -0,0 +1,125 @@
|
||||
---
|
||||
title: CrewAI
|
||||
---
|
||||
|
||||
# CrewAI
|
||||
|
||||
[CrewAI](https://www.crewai.com) is a framework for orchestrating role-playing, autonomous AI agents. By leveraging collaborative intelligence, CrewAI allows agents to work together seamlessly, tackling complex tasks.
|
||||
|
||||
The framework has a sophisticated memory system designed to significantly enhance the capabilities of AI agents. This system aids agents to remember, reason, and learn from past interactions. You can use Qdrant to store short-term memory and entity memories of CrewAI agents.
|
||||
|
||||
- Short-Term Memory
|
||||
|
||||
Temporarily stores recent interactions and outcomes using RAG, enabling agents to recall and utilize information relevant to their current context during the current executions.
|
||||
|
||||
- Entity Memory
|
||||
|
||||
Entity Memory Captures and organizes information about entities (people, places, concepts) encountered during tasks, facilitating deeper understanding and relationship mapping. Uses RAG for storing entity information.
|
||||
|
||||
## Usage with Qdrant
|
||||
|
||||
We'll learn how to customize CrewAI's default memory storage to use Qdrant.
|
||||
|
||||
### Installation
|
||||
|
||||
First, install CrewAI and Qdrant client packages:
|
||||
|
||||
```shell
|
||||
pip install 'crewai[tools]' 'qdrant-client[fastembed]'
|
||||
```
|
||||
|
||||
### Setup a CrewAI Project
|
||||
|
||||
You can learn to set up a CrewAI project [here](https://docs.crewai.com/installation#create-a-new-crewai-project). Let's assume the project was name `mycrew`.
|
||||
|
||||
### Define the Qdrant storage
|
||||
|
||||
> src/mycrew/storage.py
|
||||
|
||||
```python
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from crewai.memory.storage.rag_storage import RAGStorage
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
|
||||
class QdrantStorage(RAGStorage):
|
||||
"""
|
||||
Extends Storage to handle embeddings for memory entries using Qdrant.
|
||||
|
||||
"""
|
||||
|
||||
def __init__(self, type, allow_reset=True, embedder_config=None, crew=None):
|
||||
super().__init__(type, allow_reset, embedder_config, crew)
|
||||
|
||||
def search(
|
||||
self,
|
||||
query: str,
|
||||
limit: int = 3,
|
||||
filter: Optional[dict] = None,
|
||||
score_threshold: float = 0,
|
||||
) -> List[Any]:
|
||||
points = self.client.query(
|
||||
self.type,
|
||||
query_text=query,
|
||||
query_filter=filter,
|
||||
limit=limit,
|
||||
score_threshold=score_threshold,
|
||||
)
|
||||
results = [
|
||||
{
|
||||
"id": point.id,
|
||||
"metadata": point.metadata,
|
||||
"context": point.document,
|
||||
"score": point.score,
|
||||
}
|
||||
for point in points
|
||||
]
|
||||
|
||||
return results
|
||||
|
||||
def reset(self) -> None:
|
||||
self.client.delete_collection(self.type)
|
||||
|
||||
def _initialize_app(self):
|
||||
self.client = QdrantClient()
|
||||
if not self.client.collection_exists(self.type):
|
||||
self.client.create_collection(
|
||||
collection_name=self.type,
|
||||
vectors_config=self.client.get_fastembed_vector_params(),
|
||||
sparse_vectors_config=self.client.get_fastembed_sparse_vector_params(),
|
||||
)
|
||||
|
||||
def save(self, value: Any, metadata: Dict[str, Any]) -> None:
|
||||
self.client.add(self.type, documents=[value], metadata=[metadata or {}])
|
||||
```
|
||||
|
||||
The `add` AND `query` methods use [FastEmbed](https://github.com/qdrant/fastembed/) to vectorize data. You can however customize it if required.
|
||||
|
||||
### Instantiate your crew
|
||||
|
||||
You can learn about setting up agents and tasks for your crew [here](https://docs.crewai.com/quickstart). We can update the instantiation of `Crew` to use our storage mechanism.
|
||||
|
||||
> src/mycrew/crew.py
|
||||
|
||||
```python
|
||||
from crewai import Crew
|
||||
from crewai.memory.entity.entity_memory import EntityMemory
|
||||
from crewai.memory.short_term.short_term_memory import ShortTermMemory
|
||||
|
||||
from mycrew.storage import QdrantStorage
|
||||
|
||||
Crew(
|
||||
# Import the agents and tasks here.
|
||||
memory=True,
|
||||
entity_memory=EntityMemory(storage=QdrantStorage("entity")),
|
||||
short_term_memory=ShortTermMemory(storage=QdrantStorage("short-term")),
|
||||
)
|
||||
```
|
||||
|
||||
You can now run your Crew workflow with `crew run`. It'll use Qdrant for memory ingestion and retrieval.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [CrewAI Documentation](https://docs.crewai.com/introduction)
|
||||
- [CrewAI Examples](https://github.com/crewAIInc/crewAI-examples)
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
title: Feast
|
||||
---
|
||||
|
||||
## Feast
|
||||
|
||||
[Feast (**Fe**ature **St**ore)](https://docs.feast.dev) is an open-source feature store that helps teams operate production ML systems at scale by allowing them to define, manage, validate, and serve features for production AI/ML.
|
||||
|
||||
Qdrant is available as a supported vectorstore in Feast to integrate in your workflows.
|
||||
|
||||
## Insatallation
|
||||
|
||||
To use the Qdrant online store, you need to install Feast with the `qdrant` extra.
|
||||
|
||||
```bash
|
||||
pip install 'feast[qdrant]'
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
An example config with Qdrant could look like:
|
||||
|
||||
```yaml
|
||||
project: my_feature_repo
|
||||
registry: data/registry.db
|
||||
provider: local
|
||||
online_store:
|
||||
type: qdrant
|
||||
host: xyz-example.eu-central.aws.cloud.qdrant.io
|
||||
port: 6333
|
||||
api_key: <your-own-key>
|
||||
vector_len: 384
|
||||
# Reference: https://qdrant.tech/documentation/concepts/vectors/#named-vectors
|
||||
# vector_name: text-vec
|
||||
write_batch_size: 100
|
||||
```
|
||||
|
||||
You can refer to the Feast [reference](https://rtd.feast.dev/en/master/index.html#) for the full list of configuration options.
|
||||
|
||||
## Retrieving Documents
|
||||
|
||||
The Qdrant online store supports retrieving document vectors for a given list of entity keys. The document vectors are returned as a dictionary where the key is the entity key and the value being the vector.
|
||||
|
||||
```python
|
||||
from feast import FeatureStore
|
||||
|
||||
feature_store = FeatureStore(repo_path="feature_store.yaml")
|
||||
|
||||
query_vector = [1.0, 2.0, 3.0, 4.0, 5.0]
|
||||
top_k = 5
|
||||
|
||||
feature_values = feature_store.retrieve_online_documents(
|
||||
feature="my_feature",
|
||||
query=query_vector,
|
||||
top_k=top_k
|
||||
)
|
||||
```
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- [Feast Docs](http://docs.feast.dev/)
|
||||
- [Feast Reference](https://rtd.feast.dev/en/master/index.html/)
|
||||
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/)
|
||||
@@ -0,0 +1,78 @@
|
||||
---
|
||||
title: AWS Lakechain
|
||||
---
|
||||
|
||||
# AWS Lakechain
|
||||
|
||||
[Project Lakechain](https://awslabs.github.io/project-lakechain/) is a framework based on the AWS Cloud Development Kit (CDK), allowing to express and deploy scalable document processing pipelines on AWS using infrastructure-as-code. It emphasizes on modularity and extensibility of pipelines, and provides 60+ ready to use components for prototyping complex processing pipelines that scale out of the box to millions of documents.
|
||||
|
||||
The Qdrant storage connector available with Lakechain enables uploading vector embeddings produced by other middlewares to a Qdrant collection.
|
||||
|
||||
<aside role="status">You can find an end-to-end example usage of the Qdrant Lakechain connector <a a target="_blank" href="https://github.com/awslabs/project-lakechain/tree/main/examples/simple-pipelines/embedding-pipelines/bedrock-qdrant-pipeline">here.</a></aside>
|
||||
|
||||
To use the Qdrant storage connector, you import it in your CDK stack, and connect it to a data source providing document embeddings.
|
||||
|
||||
> You need to specify a Qdrant API key to the connector, by specifying a reference to an [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/) secret containing the API key.
|
||||
|
||||
```typescript
|
||||
import { QdrantStorageConnector } from '@project-lakechain/qdrant-storage-connector';
|
||||
import { CacheStorage } from '@project-lakechain/core';
|
||||
|
||||
class Stack extends cdk.Stack {
|
||||
constructor(scope: cdk.Construct, id: string) {
|
||||
const cache = new CacheStorage(this, 'Cache');
|
||||
|
||||
const qdrantApiKey = secrets.Secret.fromSecretNameV2(
|
||||
this,
|
||||
'QdrantApiKey',
|
||||
process.env.QDRANT_API_KEY_SECRET_NAME as string
|
||||
);
|
||||
|
||||
const connector = new QdrantStorageConnector.Builder()
|
||||
.withScope(this)
|
||||
.withIdentifier('QdrantStorageConnector')
|
||||
.withCacheStorage(cache)
|
||||
.withSource(source) // 👈 Specify a data source
|
||||
.withApiKey(qdrantApiKey)
|
||||
.withCollectionName('{collection_name}')
|
||||
.withUrl('https://xyz-example.eu-central.aws.cloud.qdrant.io:6333')
|
||||
.build();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
When the document being processed is a text document, you can choose to store the text of the document in the Qdrant payload. To do so, you can use the `withStoreText` and `withTextKey` options. If the document is not a text, this option is ignored.
|
||||
|
||||
```typescript
|
||||
const connector = new QdrantStorageConnector.Builder()
|
||||
.withScope(this)
|
||||
.withIdentifier('QdrantStorageConnector')
|
||||
.withCacheStorage(cache)
|
||||
.withSource(source)
|
||||
.withApiKey(qdrantApiKey)
|
||||
.withCollectionName('{collection_name}')
|
||||
.withStoreText(true)
|
||||
.withTextKey('my-content')
|
||||
.withUrl('https://xyz-example.eu-central.aws.cloud.qdrant.io:6333')
|
||||
.build();
|
||||
```
|
||||
|
||||
Since Qdrant supports [multiple vectors](/documentation/concepts/vectors/#named-vectors) per point, you can use the `withVectorName` option to specify one. The connector defaults to unnamed (default) vector.
|
||||
|
||||
```typescript
|
||||
const connector = new QdrantStorageConnector.Builder()
|
||||
.withScope(this)
|
||||
.withIdentifier('QdrantStorageConnector')
|
||||
.withCacheStorage(cache)
|
||||
.withSource(source)
|
||||
.withApiKey(qdrantApiKey)
|
||||
.withCollectionName('collection_name')
|
||||
.withVectorName('my-vector-name')
|
||||
.withUrl('https://xyz-example.eu-central.aws.cloud.qdrant.io:6333')
|
||||
.build();
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Introduction to Lakechain](https://awslabs.github.io/project-lakechain/general/introduction/)
|
||||
- [Lakechain Examples](https://github.com/awslabs/project-lakechain/tree/main/examples)
|
||||
@@ -54,7 +54,7 @@ func main() {
|
||||
}
|
||||
|
||||
store, err: = qdrant.New(
|
||||
qdrant.WithURL( * url),
|
||||
qdrant.WithURL(*url),
|
||||
qdrant.WithCollectionName("YOUR_COLLECTION_NAME"),
|
||||
qdrant.WithEmbedder(e),
|
||||
)
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
---
|
||||
title: LangGraph
|
||||
aliases: [ ../integrations/autogen/ ]
|
||||
---
|
||||
|
||||
# LangGraph
|
||||
|
||||
[LangGraph](https://github.com/langchain-ai/langgraph) is a library for building stateful, multi-actor applications, ideal for creating agentic workflows. It provides fine-grained control over both the flow and state of your application, crucial for creating reliable agents.
|
||||
|
||||
You can define flows that involve cycles, essential for most agentic architectures, differentiating it from DAG-based solutions. Additionally, LangGraph includes built-in persistence, enabling advanced human-in-the-loop and memory features.
|
||||
|
||||
LangGraph works seamlessly with all the components of LangChain. This means we can utilize Qdrant's [Langchain integration](/documentation/frameworks/langchain/) to create retrieval nodes in LangGraph, available in both Python and Javascript!
|
||||
|
||||
## Usage
|
||||
|
||||
- Install the required dependencies
|
||||
|
||||
```python
|
||||
$ pip install langgraph langchain_community langchain_qdrant fastembed
|
||||
```
|
||||
|
||||
```typescript
|
||||
$ npm install @langchain/langgraph langchain @langchain/qdrant @langchain/openai
|
||||
```
|
||||
|
||||
- Create a retriever tool to add to the LangGraph workflow.
|
||||
|
||||
```python
|
||||
|
||||
from langchain.tools.retriever import create_retriever_tool
|
||||
from langchain_community.embeddings import FastEmbedEmbeddings
|
||||
|
||||
from langchain_qdrant import FastEmbedSparse, QdrantVectorStore, RetrievalMode
|
||||
|
||||
# We'll set up Qdrant to retrieve documents using Hybrid search.
|
||||
# Learn more at https://qdrant.tech/articles/hybrid-search/
|
||||
retriever = QdrantVectorStore.from_texts(
|
||||
url="http://localhost:6333/",
|
||||
collection_name="langgraph-collection",
|
||||
embedding=FastEmbedEmbeddings(model_name="BAAI/bge-small-en-v1.5"),
|
||||
sparse_embedding=FastEmbedSparse(model_name="Qdrant/bm25"),
|
||||
retrieval_mode=RetrievalMode.HYBRID,
|
||||
texts=["<SOME_KNOWLEDGE_TEXT>", "<SOME_OTHER_TEXT>", ...]
|
||||
).as_retriever()
|
||||
|
||||
retriever_tool = create_retriever_tool(
|
||||
retriever,
|
||||
"retrieve_my_texts",
|
||||
"Retrieve texts stored in the Qdrant collection",
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantVectorStore } from "@langchain/qdrant";
|
||||
import { OpenAIEmbeddings } from "@langchain/openai";
|
||||
import { createRetrieverTool } from "langchain/tools/retriever";
|
||||
|
||||
const vectorStore = await QdrantVectorStore.fromTexts(
|
||||
["<SOME_KNOWLEDGE_TEXT>", "<SOME_OTHER_TEXT>"],
|
||||
[{ id: 2 }, { id: 1 }],
|
||||
new OpenAIEmbeddings(),
|
||||
{
|
||||
url: "http://localhost:6333/",
|
||||
collectionName: "goldel_escher_bach",
|
||||
}
|
||||
);
|
||||
|
||||
const retriever = vectorStore.asRetriever();
|
||||
|
||||
const tool = createRetrieverTool(
|
||||
retriever,
|
||||
{
|
||||
name: "retrieve_my_texts",
|
||||
description:
|
||||
"Retrieve texts stored in the Qdrant collection",
|
||||
},
|
||||
);
|
||||
```
|
||||
|
||||
- Add the retriever tool as a node in LangGraph
|
||||
|
||||
```python
|
||||
from langgraph.graph import StateGraph
|
||||
from langgraph.prebuilt import ToolNode
|
||||
|
||||
workflow = StateGraph()
|
||||
|
||||
# Define other the nodes which we'll cycle between.
|
||||
workflow.add_node("retrieve_qdrant", ToolNode([retriever_tool]))
|
||||
|
||||
graph = workflow.compile()
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { StateGraph } from "@langchain/langgraph";
|
||||
import { ToolNode } from "@langchain/langgraph/prebuilt";
|
||||
|
||||
// Define the graph
|
||||
const workflow = new StateGraph(SomeGraphState)
|
||||
// Define the nodes which we'll cycle between.
|
||||
.addNode("retrieve", new ToolNode([tool]));
|
||||
|
||||
const graph = workflow.compile();
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [LangGraph Documentation](https://langchain-ai.github.io/langgraph/)
|
||||
- [LangGraph End-to-End Guides](https://langchain-ai.github.io/langgraph/tutorials/)
|
||||
@@ -0,0 +1,69 @@
|
||||
---
|
||||
title: Neo4j GraphRAG
|
||||
---
|
||||
|
||||
# Neo4j GraphRAG
|
||||
|
||||
[Neo4j GraphRAG](https://neo4j.com/docs/neo4j-graphrag-python/current/) is a Python package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. As a first-party library, it offers a robust, feature-rich, and high-performance solution, with the added assurance of long-term support and maintenance directly from Neo4j. It offers a Qdrant retriever natively to search for vectors stored in a Qdrant collection.
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
pip install neo4j-graphrag[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
A vector query with Neo4j and Qdrant could look like:
|
||||
|
||||
```python
|
||||
from neo4j import GraphDatabase
|
||||
from neo4j_graphrag.retrievers import QdrantNeo4jRetriever
|
||||
from qdrant_client import QdrantClient
|
||||
from examples.embedding_biology import EMBEDDING_BIOLOGY
|
||||
|
||||
NEO4J_URL = "neo4j://localhost:7687"
|
||||
NEO4J_AUTH = ("neo4j", "password")
|
||||
|
||||
with GraphDatabase.driver(NEO4J_URL, auth=NEO4J_AUTH) as neo4j_driver:
|
||||
retriever = QdrantNeo4jRetriever(
|
||||
driver=neo4j_driver,
|
||||
client=QdrantClient(url="http://localhost:6333"),
|
||||
collection_name="{collection_name}",
|
||||
id_property_external="neo4j_id",
|
||||
id_property_neo4j="id",
|
||||
)
|
||||
|
||||
retriever.search(query_vector=[0.5523, 0.523, 0.132, 0.523, ...], top_k=5)
|
||||
```
|
||||
|
||||
Alternatively, you can use any [Langchain embeddings providers](https://python.langchain.com/docs/integrations/text_embedding/), to vectorize text queries automatically.
|
||||
|
||||
```python
|
||||
from langchain_huggingface.embeddings import HuggingFaceEmbeddings
|
||||
from neo4j import GraphDatabase
|
||||
from neo4j_graphrag.retrievers import QdrantNeo4jRetriever
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
NEO4J_URL = "neo4j://localhost:7687"
|
||||
NEO4J_AUTH = ("neo4j", "password")
|
||||
|
||||
with GraphDatabase.driver(NEO4J_URL, auth=NEO4J_AUTH) as neo4j_driver:
|
||||
embedder = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
|
||||
retriever = QdrantNeo4jRetriever(
|
||||
driver=neo4j_driver,
|
||||
client=QdrantClient(url="http://localhost:6333"),
|
||||
collection_name="{collection_name}",
|
||||
id_property_external="neo4j_id",
|
||||
id_property_neo4j="id",
|
||||
embedder=embedder,
|
||||
)
|
||||
|
||||
retriever.search(query_text="my user query", top_k=10)
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Neo4j GraphRAG Reference](https://neo4j.com/docs/neo4j-graphrag-python/current/index.html)
|
||||
- [Qdrant Retriever Reference](https://neo4j.com/docs/neo4j-graphrag-python/current/user_guide_rag.html#qdrant-neo4j-retriever-user-guide)
|
||||
- [Source](https://github.com/neo4j/neo4j-graphrag-python/tree/main/src/neo4j_graphrag/retrievers/external/qdrant)
|
||||
@@ -0,0 +1,83 @@
|
||||
---
|
||||
title: Ragbits
|
||||
---
|
||||
|
||||
# Ragbits
|
||||
|
||||
[Ragbit](https://ragbits.deepsense.ai) is a Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. It prioritizes developer experience by providing a simple and intuitive API. It also includes a comprehensive set of tools for seamlessly building, testing, and deploying your RAG applications efficiently.
|
||||
|
||||
Qdrant is available as a vectorstore in Ragbits to ingest and search search documents from a collection.
|
||||
|
||||
## Installation
|
||||
|
||||
Install the Python package that comes bundled with the Qdrant integration.
|
||||
|
||||
```bash
|
||||
pip install ragbits
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
An example usage of Ragbits and Qdrant would look something like this:
|
||||
|
||||
The following example uses [OpenAI embeddings](https://platform.openai.com/docs/guides/embeddings) via [LiteLLM](https://www.litellm.ai).
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
from qdrant_client import AsyncQdrantClient
|
||||
|
||||
from ragbits.core.embeddings.litellm import LiteLLMEmbeddings
|
||||
from ragbits.core.vector_stores.qdrant import QdrantVectorStore
|
||||
from ragbits.document_search import DocumentSearch, SearchConfig
|
||||
from ragbits.document_search.documents.document import DocumentMeta
|
||||
|
||||
documents = [
|
||||
DocumentMeta.create_text_document_from_literal(
|
||||
"RIP boiled water. You will be mist."
|
||||
),
|
||||
DocumentMeta.create_text_document_from_literal(
|
||||
"Why programmers don't like to swim? Because they're scared of the floating points."
|
||||
),
|
||||
DocumentMeta.create_text_document_from_literal("This one is completely unrelated."),
|
||||
]
|
||||
|
||||
|
||||
async def main() -> None:
|
||||
embedder = LiteLLMEmbeddings(
|
||||
model="text-embedding-3-small",
|
||||
)
|
||||
vector_store = QdrantVectorStore(
|
||||
client=AsyncQdrantClient(url="http://localhost:6333"),
|
||||
collection_name="{collection_name}",
|
||||
)
|
||||
document_search = DocumentSearch(
|
||||
embedder=embedder,
|
||||
vector_store=vector_store,
|
||||
)
|
||||
|
||||
await document_search.ingest(documents)
|
||||
|
||||
all_documents = await vector_store.list()
|
||||
print([doc.metadata["content"] for doc in all_documents])
|
||||
|
||||
query = "I write computer software. Tell me something."
|
||||
vector_store_kwargs = {
|
||||
"k": 1,
|
||||
"max_distance": None,
|
||||
}
|
||||
results = await document_search.search(
|
||||
query,
|
||||
config=SearchConfig(vector_store_kwargs=vector_store_kwargs),
|
||||
)
|
||||
|
||||
print(f"Documents similar to: {query}")
|
||||
print([element.get_key() for element in results])
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- Ragbits [Documentation](http://ragbits.deepsense.ai)
|
||||
- [Source Code](https://github.com/deepsense-ai/ragbits)
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
title: Rig-rs
|
||||
---
|
||||
|
||||
# Rig-rs
|
||||
|
||||
[Rig](http://rig.rs) is a Rust library for building scalable, modular, and ergonomic LLM-powered applications. It has full support for LLM completion and embedding workflows with minimal boiler plate.
|
||||
|
||||
Rig supports Qdrant as a vectorstore to ingest and search for documents semantically.
|
||||
|
||||
## Installation
|
||||
|
||||
```console
|
||||
cargo add rig-core rig-qdrant qdrant-client
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Here's an example ingest and retrieve flow using Rig and Qdrant.
|
||||
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
qdrant::{PointStruct, QueryPointsBuilder, UpsertPointsBuilder},
|
||||
Payload, Qdrant,
|
||||
};
|
||||
use rig::{
|
||||
embeddings::EmbeddingsBuilder,
|
||||
providers::openai::{Client, TEXT_EMBEDDING_3_SMALL},
|
||||
vector_store::VectorStoreIndex,
|
||||
};
|
||||
use rig_qdrant::QdrantVectorStore;
|
||||
use serde_json::json;
|
||||
|
||||
const COLLECTION_NAME: &str = "rig-collection";
|
||||
|
||||
// Initialize Qdrant client.
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
// Initialize OpenAI client.
|
||||
let openai_client = Client::new("<OPENAI_API_KEY>");
|
||||
let model = openai_client.embedding_model(TEXT_EMBEDDING_3_SMALL);
|
||||
|
||||
let documents = EmbeddingsBuilder::new(model.clone())
|
||||
.simple_document("0981d983-a5f8-49eb-89ea-f7d3b2196d2e", "Definition of a *flurbo*: A flurbo is a green alien that lives on cold planets")
|
||||
.simple_document("62a36d43-80b6-4fd6-990c-f75bb02287d1", "Definition of a *glarb-glarb*: A glarb-glarb is a ancient tool used by the ancestors of the inhabitants of planet Jiro to farm the land.")
|
||||
.simple_document("f9e17d59-32e5-440c-be02-b2759a654824", "Definition of a *linglingdong*: A term used by inhabitants of the far side of the moon to describe humans.")
|
||||
.build()
|
||||
.await?;
|
||||
|
||||
let points: Vec<PointStruct> = documents
|
||||
.into_iter()
|
||||
.map(|d| {
|
||||
let vec: Vec<f32> = d.embeddings[0].vec.iter().map(|&x| x as f32).collect();
|
||||
PointStruct::new(
|
||||
d.id,
|
||||
vec,
|
||||
Payload::try_from(json!({
|
||||
"document": d.document,
|
||||
}))
|
||||
.unwrap(),
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
|
||||
client
|
||||
.upsert_points(UpsertPointsBuilder::new(COLLECTION_NAME, points))
|
||||
.await?;
|
||||
|
||||
let query_params = QueryPointsBuilder::new(COLLECTION_NAME).with_payload(true);
|
||||
let vector_store = QdrantVectorStore::new(client, model, query_params.build());
|
||||
|
||||
let results = vector_store
|
||||
.top_n::<serde_json::Value>("Define a glarb-glarb?", 1)
|
||||
.await?;
|
||||
|
||||
println!("Results: {:?}", results);
|
||||
```
|
||||
|
||||
## Further reading
|
||||
|
||||
- [Rig-rs Documentation](https://rig.rs)
|
||||
- [Source Code](https://github.com/0xPlaygrounds/rig)
|
||||
@@ -0,0 +1,171 @@
|
||||
---
|
||||
title: OpenAI Swarm
|
||||
---
|
||||
|
||||
# Swarm
|
||||
|
||||
[OpenAI Swarm](https://github.com/openai/swarm) is a Python framework for managing multiple AI agents that can work together. Instead of relying on a single LLM instance to perform all tasks, Swarm allows you to build specialized agents that communicate and collaborate, like a team of experts with unique skills.
|
||||
|
||||
## Getting Started
|
||||
|
||||
To start using Swarm, follow these steps:
|
||||
|
||||
- Install Swarm
|
||||
|
||||
```bash
|
||||
pip install git+https://github.com/openai/swarm.git
|
||||
```
|
||||
|
||||
- Set up your OpenAI API key
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY="<YOUR_KEY>"
|
||||
```
|
||||
|
||||
## How Swarm Works
|
||||
|
||||
In Swarm, agents represent individual team members with specific roles and instructions. Each agent can execute tasks or hand off the conversation to another agent, depending on the situation.
|
||||
|
||||
An `Agent` instance simply encapsulates a set of instructions with a set of functions, and has the capability to hand off execution to another Agent.
|
||||
|
||||
These simple building blocks allow you to create complex workflows with a simple mental model.
|
||||
|
||||
## Creating Your First Agents
|
||||
|
||||
Here’s a basic example of two agents:
|
||||
|
||||
- **Agent A**: A helpful assistant.
|
||||
- **Agent B**: An arithmetric specialist.
|
||||
|
||||
Agent A transfers the conversation to Agent B when requested.
|
||||
|
||||
```python
|
||||
from swarm import Swarm, Agent
|
||||
|
||||
client = Swarm()
|
||||
|
||||
# Define Agent B
|
||||
agent_b = Agent(
|
||||
name="Agent B",
|
||||
instructions="Arithmetic solving expertise holder.",
|
||||
)
|
||||
|
||||
def transfer_to_agent_b():
|
||||
return agent_b
|
||||
|
||||
# Define Agent A
|
||||
agent_a = Agent(
|
||||
name="Agent A",
|
||||
instructions="You are a helpful agent.",
|
||||
functions=[transfer_to_agent_b],
|
||||
)
|
||||
|
||||
# Run the interaction
|
||||
response = client.run(
|
||||
agent=agent_a,
|
||||
messages=[{"role": "user", "content": "I want some help with numbers."}],
|
||||
)
|
||||
print(response)
|
||||
```
|
||||
|
||||
In this example, Agent A passes the conversation to Agent B when deemed necessary.
|
||||
|
||||
## Features of an agent
|
||||
|
||||
### 1. **Instructions**
|
||||
|
||||
Agent instructions define their behavior. These are translated into system prompts for conversations. Only the active agent's instructions are used during an interaction.
|
||||
|
||||
### 2. **Functions**
|
||||
|
||||
Agents can execute Python functions, enabling them to perform tasks like processing data or querying databases. Swarm automatically converts functions into a JSON Schema that is passed into Chat Completions tools.
|
||||
|
||||
Example:
|
||||
|
||||
```python
|
||||
def greet(context_variables, language):
|
||||
user_name = context_variables["user_name"]
|
||||
greeting = "Hola" if language.lower() == "spanish" else "Hello"
|
||||
print(f"{greeting}, {user_name}!")
|
||||
return "Done"
|
||||
|
||||
agent = Agent(
|
||||
name="Greeter Agent",
|
||||
functions=[greet],
|
||||
)
|
||||
|
||||
client.run(
|
||||
agent=agent,
|
||||
messages=[{"role": "user", "content": "Greet me in Spanish."}],
|
||||
context_variables={"user_name": "John"},
|
||||
)
|
||||
```
|
||||
|
||||
Errors are handled gracefully by appending an error response to the conversation.
|
||||
|
||||
### 3. **Handoffs**
|
||||
|
||||
If a function returns another agent, the system transfers control to that agent.
|
||||
|
||||
## Integrating Swarm with Qdrant
|
||||
|
||||
You can connect you Swarm agents to retrieve or ingest data into a Qdrant collection. Thereby building you knowledge base. Here’s how to enable an agent to retrieve information from Qdrant.
|
||||
|
||||
Assume you have a Qdrant [collection created](https://qdrant.tech/documentation/concepts/collections/#create-a-collection) using the `"text-embedding-3-small"` model. The payload structure includes a `text` field for knowledge storage.
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
from openai import OpenAI
|
||||
|
||||
# Initialize clients
|
||||
openai_client = OpenAI()
|
||||
qdrant = qdrant_client.QdrantClient(host="localhost")
|
||||
|
||||
# Configuration
|
||||
EMBEDDING_MODEL = "text-embedding-3-small"
|
||||
COLLECTION_NAME = "help_center"
|
||||
LIMIT = 5
|
||||
SCORE_THRESHOLD = 0.7
|
||||
|
||||
# Function to query Qdrant
|
||||
def query_qdrant(query):
|
||||
"""Retrieve semantically relevant content from Qdrant."""
|
||||
embedded_query = openai_client.embeddings.create(
|
||||
input=query,
|
||||
model=EMBEDDING_MODEL,
|
||||
).data[0].embedding
|
||||
|
||||
results = qdrant.query_points(
|
||||
collection_name=COLLECTION_NAME,
|
||||
query=embedded_query,
|
||||
limit=LIMIT,
|
||||
score_threshold=SCORE_THRESHOLD,
|
||||
).points
|
||||
|
||||
if results:
|
||||
return {"response": "\n".join([point.payload["text"] for point in results])}
|
||||
else:
|
||||
return {"response": "No results found."}
|
||||
|
||||
# Define agents
|
||||
qdrant_agent = Agent(
|
||||
name="Qdrant Agent",
|
||||
instructions="Retrieve relevant info from a knowledge base stored in Qdrant.",
|
||||
functions=[query_qdrant],
|
||||
)
|
||||
|
||||
def transfer_to_qdrant():
|
||||
return qdrant_agent
|
||||
|
||||
main_agent = Agent(
|
||||
name="Main Agent",
|
||||
instructions="Handle user queries and delegate searches to Qdrant.",
|
||||
functions=[transfer_to_qdrant],
|
||||
)
|
||||
```
|
||||
|
||||
Our `qdrant_agent` can now query a Qdrant collection whenever deemed necessary to answer a user query.
|
||||
|
||||
## Further Reading
|
||||
|
||||
You can find more usage examples in the [Swarm repo](https://github.com/openai/swarm/blob/main/examples/) that further describe its capabilities.
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
title: Sycamore
|
||||
---
|
||||
|
||||
## Sycamore
|
||||
|
||||
[Sycamore](https://sycamore.readthedocs.io/en/stable/) is an LLM-powered data preparation, processing, and analytics system for complex, unstructured documents like PDFs, HTML, presentations, and more. With Aryn, you can prepare data for GenAI and RAG applications, power high-quality document processing workflows, and run analytics on large document collections with natural language.
|
||||
|
||||
You can use the Qdrant connector to write into and read documents from Qdrant collections.
|
||||
|
||||
<aside role="status">You can find an end-to-end example usage of the Qdrant connector <a a target="_blank" href="https://github.com/aryn-ai/sycamore/blob/main/examples/simple_qdrant.py">here.</a></aside>
|
||||
|
||||
## Writing to Qdrant
|
||||
|
||||
To write a Docset to a Qdrant collection in Sycamore, use the `docset.write.qdrant(....)` function. The Qdrant writer accepts the following arguments:
|
||||
|
||||
- `client_params`: Parameters that are passed to the Qdrant client constructor. See more information in the [Client API Reference](https://python-client.qdrant.tech/qdrant_client.qdrant_client).
|
||||
- `collection_params`: Parameters that are passed into the `qdrant_client.QdrantClient.create_collection` method. See more information in the [Client API Reference](https://python-client.qdrant.tech/_modules/qdrant_client/qdrant_client#QdrantClient.create_collection).
|
||||
- `vector_name`: The name of the vector in the Qdrant collection. Defaults to `None`.
|
||||
- `execute`: Execute the pipeline and write to Qdrant on adding this operator. If `False`, will return a `DocSet` with this write in the plan. Defaults to `True`.
|
||||
- `kwargs`: Keyword arguments to pass to the underlying execution engine.
|
||||
|
||||
```python
|
||||
ds.write.qdrant(
|
||||
{
|
||||
"url": "http://localhost:6333",
|
||||
"timeout": 50,
|
||||
},
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"vectors_config": {
|
||||
"size": 384,
|
||||
"distance": "Cosine",
|
||||
},
|
||||
},
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
## Reading from Qdrant
|
||||
|
||||
To read a Docset from a Qdrant collection in Sycamore, use the `docset.read.qdrant(....)` function. The Qdrant reader accepts the following arguments:
|
||||
|
||||
- `client_params`: Parameters that are passed to the Qdrant client constructor. See more information in the[Client API Reference](https://python-client.qdrant.tech/qdrant_client.qdrant_client).
|
||||
- `query_params`: Parameters that are passed into the `qdrant_client.QdrantClient.query_points` method. See more information in the [Client API Reference](https://python-client.qdrant.tech/_modules/qdrant_client/qdrant_client#QdrantClient.query_points).
|
||||
- `kwargs`: Keyword arguments to pass to the underlying execution engine.
|
||||
|
||||
```python
|
||||
docs = ctx.read.qdrant(
|
||||
{
|
||||
"url": "https://xyz-example.eu-central.aws.cloud.qdrant.io:6333",
|
||||
"api_key": "<paste-your-api-key-here>",
|
||||
},
|
||||
{"collection_name": "{collection_name}", "limit": 100, "using": "{optional_vector_name}"},
|
||||
).take_all()
|
||||
|
||||
```
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- [Sycamore Reference](https://sycamore.readthedocs.io/en/stable/)
|
||||
- [Sycamore](https://github.com/aryn-ai/sycamore/tree/main/examples)
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: Testcontainers
|
||||
aliases: [ ../frameworks/testcontainers/ ]
|
||||
aliases: [ ../infrastructure/testcontainers/ ]
|
||||
---
|
||||
|
||||
# Testcontainers
|
||||
@@ -3,4 +3,5 @@ title: Guides
|
||||
weight: 9
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: true
|
||||
partition: qdrant
|
||||
---
|
||||
@@ -13,6 +13,7 @@ Qdrant exposes administration tools which enable to modify at runtime the behavi
|
||||
|
||||
A locking API enables users to restrict the possible operations on a qdrant process.
|
||||
It is important to mention that:
|
||||
|
||||
- The configuration is not persistent therefore it is necessary to lock again following a restart.
|
||||
- Locking applies to a single node only. It is necessary to call lock on all the desired nodes in a distributed deployment setup.
|
||||
|
||||
|
||||
@@ -8,85 +8,120 @@ aliases:
|
||||
|
||||
# Configuration
|
||||
|
||||
To change or correct Qdrant's behavior, default collection settings, and network interface parameters, you can use configuration files.
|
||||
Qdrant ships with sensible defaults for collection and network settings that are suitable for most use cases. You can view these defaults in the [Qdrant source](https://github.com/qdrant/qdrant/blob/master/config/config.yaml). If you need to customize the settings, you can do so using configuration files and environment variables.
|
||||
|
||||
The default configuration file is located at [config/config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml).
|
||||
<aside role="status">
|
||||
Qdrant Cloud does not allow modifying the Qdrant configuration.
|
||||
</aside>
|
||||
|
||||
To change the default configuration, add a new configuration file and specify
|
||||
the path with `--config-path path/to/custom_config.yaml`. If running in
|
||||
production mode, you could also choose to overwrite `config/production.yaml`.
|
||||
See [ordering](#order-and-priority) for details on how configurations are
|
||||
loaded.
|
||||
## Configuration Files
|
||||
|
||||
The [Installation](/documentation/guides/installation/) guide contains examples of how to set up Qdrant with a custom configuration for the different deployment methods.
|
||||
To customize Qdrant, you can mount your configuration file in any of the following locations. This guide uses `.yaml` files, but Qdrant also supports other formats such as `.toml`, `.json`, and `.ini`.
|
||||
|
||||
## Order and priority
|
||||
1. **Main Configuration: `qdrant/config/config.yaml`**
|
||||
|
||||
*Effective as of v1.2.1*
|
||||
Mount your custom `config.yaml` file to override default settings:
|
||||
|
||||
Multiple configurations may be loaded on startup. All of them are merged into a
|
||||
single effective configuration that is used by Qdrant.
|
||||
```bash
|
||||
docker run -p 6333:6333 \
|
||||
-v $(pwd)/config.yaml:/qdrant/config/config.yaml \
|
||||
qdrant/qdrant
|
||||
```
|
||||
|
||||
Configurations are loaded in the following order, if present:
|
||||
2. **Environment-Specific Configuration: `config/{RUN_MODE}.yaml`**
|
||||
|
||||
1. Embedded base configuration ([source](https://github.com/qdrant/qdrant/blob/master/config/config.yaml))
|
||||
2. File `config/config.yaml`
|
||||
3. File `config/{RUN_MODE}.yaml` (such as `config/production.yaml`)
|
||||
4. File `config/local.yaml`
|
||||
5. Config provided with `--config-path PATH` (if set)
|
||||
6. [Environment variables](#environment-variables)
|
||||
Qdrant looks for an environment-specific configuration file based on the `RUN_MODE` variable. By default, the [official Docker image](https://hub.docker.com/r/qdrant/qdrant) uses `RUN_MODE=production`, meaning it will look for `config/production.yaml`.
|
||||
|
||||
This list is from least to most significant. Properties in later configurations
|
||||
will overwrite those loaded before it. For example, a property set with
|
||||
`--config-path` will overwrite those in other files.
|
||||
You can override this by setting `RUN_MODE` to another value (e.g., `dev`), and providing the corresponding file:
|
||||
|
||||
Most of these files are included by default in the Docker container. But it is
|
||||
likely that they are absent on your local machine if you run the `qdrant` binary
|
||||
manually.
|
||||
```bash
|
||||
docker run -p 6333:6333 \
|
||||
-v $(pwd)/dev.yaml:/qdrant/config/dev.yaml \
|
||||
-e RUN_MODE=dev \
|
||||
qdrant/qdrant
|
||||
```
|
||||
|
||||
If file 2 or 3 are not found, a warning is shown on startup.
|
||||
If file 5 is provided but not found, an error is shown on startup.
|
||||
3. **Local Configuration: `config/local.yaml`**
|
||||
|
||||
Other supported configuration file formats and extensions include: `.toml`, `.json`, `.ini`.
|
||||
The `local.yaml` file is typically used for machine-specific settings that are not tracked in version control:
|
||||
|
||||
## Environment variables
|
||||
```bash
|
||||
docker run -p 6333:6333 \
|
||||
-v $(pwd)/local.yaml:/qdrant/config/local.yaml \
|
||||
qdrant/qdrant
|
||||
```
|
||||
|
||||
It is possible to set configuration properties using environment variables.
|
||||
Environment variables are always the most significant and cannot be overwritten
|
||||
(see [ordering](#order-and-priority)).
|
||||
4. **Custom Configuration via `--config-path`**
|
||||
|
||||
All environment variables are prefixed with `QDRANT__` and are separated with
|
||||
`__`.
|
||||
You can specify a custom configuration file path using the `--config-path` argument. This will override other configuration files:
|
||||
|
||||
These variables:
|
||||
```bash
|
||||
docker run -p 6333:6333 \
|
||||
-v $(pwd)/config.yaml:/path/to/config.yaml \
|
||||
qdrant/qdrant \
|
||||
./qdrant --config-path /path/to/config.yaml
|
||||
```
|
||||
|
||||
For details on how these configurations are loaded and merged, see the [loading order and priority](#loading-order-and-priority). The full list of available configuration options can be found [below](#configuration-options).
|
||||
|
||||
## Environment Variables
|
||||
|
||||
You can also configure Qdrant using environment variables, which always take the highest priority and override any file-based settings.
|
||||
|
||||
Environment variables follow this format: they should be prefixed with `QDRANT__`, and nested properties should be separated by double underscores (`__`). For example:
|
||||
|
||||
```bash
|
||||
QDRANT__LOG_LEVEL=INFO
|
||||
QDRANT__SERVICE__HTTP_PORT=6333
|
||||
QDRANT__SERVICE__ENABLE_TLS=1
|
||||
QDRANT__TLS__CERT=./tls/cert.pem
|
||||
QDRANT__TLS__CERT_TTL=3600
|
||||
docker run -p 6333:6333 \
|
||||
-e QDRANT__LOG_LEVEL=INFO \
|
||||
-e QDRANT__SERVICE__API_KEY=<MY_SECRET_KEY> \
|
||||
-e QDRANT__SERVICE__ENABLE_TLS=1 \
|
||||
-e QDRANT__TLS__CERT=./tls/cert.pem \
|
||||
qdrant/qdrant
|
||||
```
|
||||
|
||||
result in this configuration:
|
||||
This results in the following configuration:
|
||||
|
||||
```yaml
|
||||
log_level: INFO
|
||||
service:
|
||||
http_port: 6333
|
||||
enable_tls: true
|
||||
api_key: <MY_SECRET_KEY>
|
||||
tls:
|
||||
cert: ./tls/cert.pem
|
||||
cert_ttl: 3600
|
||||
```
|
||||
|
||||
To run Qdrant locally with a different HTTP port you could use:
|
||||
## Loading Order and Priority
|
||||
|
||||
```bash
|
||||
QDRANT__SERVICE__HTTP_PORT=1234 ./qdrant
|
||||
During startup, Qdrant merges multiple configuration sources into a single effective configuration. The loading order is as follows (from least to most significant):
|
||||
|
||||
1. Embedded default configuration
|
||||
2. `config/config.yaml`
|
||||
3. `config/{RUN_MODE}.yaml`
|
||||
4. `config/local.yaml`
|
||||
5. Custom configuration file
|
||||
6. Environment variables
|
||||
|
||||
### Overriding Behavior
|
||||
|
||||
Settings from later sources in the list override those from earlier sources:
|
||||
|
||||
- Settings in `config/{RUN_MODE}.yaml` (3) will override those in `config/config.yaml` (2).
|
||||
- A custom configuration file provided via `--config-path` (5) will override all other file-based settings.
|
||||
- Environment variables (6) have the highest priority and will override any settings from files.
|
||||
|
||||
## Configuration Validation
|
||||
|
||||
Qdrant validates the configuration during startup. If any issues are found, the server will terminate immediately, providing information about the error. For example:
|
||||
|
||||
```console
|
||||
Error: invalid type: 64-bit integer `-1`, expected an unsigned 64-bit or smaller integer for key `storage.hnsw_index.max_indexing_threads` in config/production.yaml
|
||||
```
|
||||
|
||||
## Configuration file example
|
||||
This ensures that misconfigurations are caught early, preventing Qdrant from running with invalid settings.
|
||||
|
||||
## Configuration Options
|
||||
|
||||
The following YAML example describes the available configuration options.
|
||||
|
||||
```yaml
|
||||
log_level: INFO
|
||||
@@ -118,10 +153,10 @@ storage:
|
||||
# endpoint_url: ""
|
||||
|
||||
# Where to store temporary files
|
||||
# If null, temporary snapshot are stored in: storage/snapshots_temp/
|
||||
# If null, temporary snapshots are stored in: storage/snapshots_temp/
|
||||
temp_path: null
|
||||
|
||||
# If true - point's payload will not be stored in memory.
|
||||
# If true - the point's payload will not be stored in memory.
|
||||
# It will be read from the disk every time it is requested.
|
||||
# This setting saves RAM by (slightly) increasing the response time.
|
||||
# Note: those payload values that are involved in filtering and are indexed - remain in RAM.
|
||||
@@ -178,12 +213,12 @@ storage:
|
||||
# Default is to allow 1 transfer.
|
||||
# If null - allow unlimited transfers.
|
||||
#outgoing_shard_transfers_limit: 1
|
||||
|
||||
|
||||
# Enable async scorer which uses io_uring when rescoring.
|
||||
# Only supported on Linux, must be enabled in your kernel.
|
||||
# See: <https://qdrant.tech/articles/io_uring/#and-what-about-qdrant>
|
||||
#async_scorer: false
|
||||
|
||||
|
||||
optimizers:
|
||||
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
|
||||
deleted_threshold: 0.2
|
||||
@@ -216,8 +251,8 @@ storage:
|
||||
# To enable memmap storage, lower the threshold
|
||||
# Note: 1Kb = 1 vector of size 256
|
||||
# To explicitly disable mmap optimization, set to `0`.
|
||||
# If not set, will be disabled by default.
|
||||
memmap_threshold_kb: null
|
||||
# If not set, will be disabled by default. Previously this was called memmap_threshold_kb.
|
||||
memmap_threshold: null
|
||||
|
||||
# Maximum size (in KiloBytes) of vectors allowed for plain index.
|
||||
# Default value based on https://github.com/google-research/google-research/blob/master/scann/docs/algorithms.md
|
||||
@@ -242,7 +277,7 @@ storage:
|
||||
# vacuum_min_vector_number: 1000
|
||||
# default_segment_number: 0
|
||||
# max_segment_size_kb: null
|
||||
# memmap_threshold_kb: null
|
||||
# memmap_threshold: null
|
||||
# indexing_threshold_kb: 20000
|
||||
# flush_interval_sec: 5
|
||||
# max_optimization_threads: null
|
||||
@@ -298,6 +333,32 @@ storage:
|
||||
# More info: https://qdrant.tech/documentation/guides/quantization
|
||||
quantization: null
|
||||
|
||||
# Default strict mode parameters for newly created collections.
|
||||
strict_mode:
|
||||
# Whether strict mode is enabled for a collection or not.
|
||||
enabled: false
|
||||
|
||||
# Max allowed `limit` parameter for all APIs that don't have their own max limit.
|
||||
max_query_limit: null
|
||||
|
||||
# Max allowed `timeout` parameter.
|
||||
max_timeout: null
|
||||
|
||||
# Allow usage of unindexed fields in retrieval based (eg. search) filters.
|
||||
unindexed_filtering_retrieve: null
|
||||
|
||||
# Allow usage of unindexed fields in filtered updates (eg. delete by payload).
|
||||
unindexed_filtering_update: null
|
||||
|
||||
# Max HNSW value allowed in search parameters.
|
||||
search_max_hnsw_ef: null
|
||||
|
||||
# Whether exact search is allowed or not.
|
||||
search_allow_exact: null
|
||||
|
||||
# Max oversampling value allowed in search.
|
||||
search_max_oversampling: null
|
||||
|
||||
service:
|
||||
# Maximum size of POST data in a single request in megabytes
|
||||
max_request_size_mb: 32
|
||||
@@ -378,12 +439,11 @@ cluster:
|
||||
# We encourage you NOT to change this parameter unless you know what you are doing.
|
||||
tick_period_ms: 100
|
||||
|
||||
|
||||
# Set to true to prevent service from sending usage statistics to the developers.
|
||||
# Read more: https://qdrant.tech/documentation/guides/telemetry
|
||||
# Defaults: false
|
||||
telemetry_disabled: false
|
||||
|
||||
|
||||
# TLS configuration.
|
||||
# Required if either service.enable_tls or cluster.p2p.enable_tls is true.
|
||||
tls:
|
||||
@@ -408,19 +468,3 @@ tls:
|
||||
# If `null` - TTL is disabled.
|
||||
cert_ttl: 3600
|
||||
```
|
||||
|
||||
## Validation
|
||||
|
||||
*Available since v1.1.1*
|
||||
|
||||
The configuration is validated on startup. If a configuration is loaded but
|
||||
validation fails, a warning is logged. E.g.:
|
||||
|
||||
```text
|
||||
WARN Settings configuration file has validation errors:
|
||||
WARN - storage.optimizers.memmap_threshold: value 123 invalid, must be 1000 or larger
|
||||
WARN - storage.hnsw_index.m: value 1 invalid, must be from 4 to 10000
|
||||
```
|
||||
|
||||
The server will continue to operate. Any validation errors should be fixed as
|
||||
soon as possible though to prevent problematic behavior.
|
||||
|
||||
@@ -535,8 +535,7 @@ client.upsert(
|
||||
```
|
||||
|
||||
```typescript
|
||||
|
||||
client.upsertPoints("{collection_name}", {
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1111,
|
||||
@@ -753,8 +752,6 @@ to recover dead shards.
|
||||
|
||||
## Replication
|
||||
|
||||
*Available as of v0.11.0*
|
||||
|
||||
Qdrant allows you to replicate shards between nodes in the cluster.
|
||||
|
||||
Shard replication increases the reliability of the cluster by keeping several copies of a shard spread across the cluster.
|
||||
@@ -762,9 +759,9 @@ This ensures the availability of the data in case of node failures, except if al
|
||||
|
||||
### Replication factor
|
||||
|
||||
When you create a collection, you can control how many shard replicas you'd like to store by changing the `replication_factor`. By default, `replication_factor` is set to "1", meaning no additional copy is maintained automatically. You can change that by setting the `replication_factor` when you create a collection.
|
||||
When you create a collection, you can control how many shard replicas you'd like to store by changing the `replication_factor`. By default, `replication_factor` is set to "1", meaning no additional copy is maintained automatically. The default can be changed in the [Qdrant configuration](/documentation/guides/configuration/#configuration-options). You can change that by setting the `replication_factor` when you create a collection.
|
||||
|
||||
Currently, the replication factor of a collection can only be configured at creation time.
|
||||
The `replication_factor` can be updated for an existing collection, but the effect of this depends on how you're running Qdrant. If you're hosting the open source version of Qdrant yourself, changing the replication factor after collection creation doesn't do anything. You can manually [create](#creating-new-shard-replicas) or drop shard replicas to achieve your desired replication factor. In Qdrant Cloud (including Hybrid Cloud, Private Cloud) your shards will automatically be replicated or dropped to match your configured replication factor.
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
@@ -894,7 +891,7 @@ Since a replication factor of "2" would require twice as much storage space, it
|
||||
|
||||
### Creating new shard replicas
|
||||
|
||||
It is possible to create or delete replicas manually on an existing collection using the [Update collection cluster setup API](https://api.qdrant.tech/master/api-reference/distributed/update-collection-cluster).
|
||||
It is possible to create or delete replicas manually on an existing collection using the [Update collection cluster setup API](https://api.qdrant.tech/master/api-reference/distributed/update-collection-cluster). This is usually only necessary if you run Qdrant open-source. In Qdrant Cloud shard replication is handled and updated automatically, matching the configured `replication_factor`.
|
||||
|
||||
A replica can be added on a specific peer by specifying the peer from which to replicate.
|
||||
|
||||
@@ -1160,6 +1157,17 @@ client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||
|
||||
Write operations will fail if the number of active replicas is less than the `write_consistency_factor`.
|
||||
|
||||
The configuration of the `write_consistency_factor` is important for adjusting the cluster's behavior when some nodes go offline due to restarts, upgrades, or failures.
|
||||
|
||||
By default, the cluster continues to accept updates as long as at least one replica of each shard is online. However, this behavior means that once an offline replica is restored, it will require additional synchronization with the rest of the cluster. In some cases, this synchronization can be resource-intensive and undesirable.
|
||||
|
||||
Setting the `write_consistency_factor` to match the replication factor modifies the cluster's behavior so that unreplicated updates are rejected, preventing the need for extra synchronization.
|
||||
|
||||
If the update is applied to enough replicas - according to the `write_consistency_factor` - the update will return a successful status. Any replicas that failed to apply the update will be temporarily disabled and are automatically recovered to keep data consistency. If the update could not be applied to enough replicas, it'll return an error and may be partially applied. The user must submit the operation again to ensure data consistency.
|
||||
|
||||
For asynchronous updates and injection pipelines capable of handling errors and retries, this strategy might be preferable.
|
||||
|
||||
|
||||
### Read consistency
|
||||
|
||||
Read `consistency` can be specified for most read requests and will ensure that the returned result
|
||||
|
||||
@@ -619,7 +619,7 @@ client.CreateFieldIndex(context.Background(), &qdrant.CreateFieldIndexCollection
|
||||
})
|
||||
```
|
||||
|
||||
`is_tenant=true` parameter is optional, but specifying it provides storage with additional inforamtion about the usage patterns the collection is going to use.
|
||||
`is_tenant=true` parameter is optional, but specifying it provides storage with additional information about the usage patterns the collection is going to use.
|
||||
When specified, storage structure will be organized in a way to co-locate vectors of the same tenant together, which can significantly improve performance in some cases.
|
||||
|
||||
|
||||
|
||||
@@ -10,6 +10,7 @@ aliases:
|
||||
Different use cases require different balances between memory usage, search speed, and precision. Qdrant is designed to be flexible and customizable so you can tune it to your specific needs.
|
||||
|
||||
This guide will walk you three main optimization strategies:
|
||||
|
||||
- High Speed Search & Low Memory Usage
|
||||
- High Precision & Low Memory Usage
|
||||
- High Precision & High Speed Search
|
||||
|
||||
@@ -35,7 +35,9 @@ service:
|
||||
Or alternatively, you can use the environment variable:
|
||||
|
||||
```bash
|
||||
export QDRANT__SERVICE__API_KEY=your_secret_api_key_here
|
||||
docker run -p 6333:6333 \
|
||||
-e QDRANT__SERVICE__API_KEY=your_secret_api_key_here \
|
||||
qdrant/qdrant
|
||||
```
|
||||
|
||||
<aside role="alert"><a href="#tls">TLS</a> must be used to prevent leaking the API key over an unencrypted connection.</aside>
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
---
|
||||
title: Hybrid Cloud
|
||||
weight: 13
|
||||
partition: cloud
|
||||
---
|
||||
|
||||
# Qdrant Hybrid Cloud
|
||||
@@ -29,6 +30,13 @@ The Qdrant Kubernetes Operator will manage the Qdrant databases within your Kube
|
||||
|
||||
Both component's access is limited to the Kubernetes namespace that you chose during the onboarding process.
|
||||
|
||||
The Cloud Agent only sends telemetry data and status information to the Qdrant Cloud platform. It does not send any user data or sensitive information. The telemetry data includes:
|
||||
|
||||
* The health status and resource (CPU, memory, disk and network) usage of the Qdrant databases and Qdrant control plane components.
|
||||
* Information about the Qdrant databases, such as the number, name and configuration of collections, the number of vectors, the number of queries, and the number of indexing operations.
|
||||
* Telemetry and notification data from the Qdrant databases.
|
||||
* Kubernetes operations and scheduling events reported for the Qdrant databases and Qdrant control plane components.
|
||||
|
||||
After the initial onboarding, the lifecycle of these components will be controlled by the Qdrant Cloud platform via the built-in Helm controller.
|
||||
|
||||
You don't need to expose your Kubernetes Cluster to the Qdrant Cloud platform, you don't need to open any ports for incoming traffic, and you don't need to provide any Kubernetes or cloud provider credentials to the Qdrant Cloud platform.
|
||||
|
||||
@@ -7,6 +7,8 @@ weight: 2
|
||||
|
||||
Once you have created a Hybrid Cloud Environment, you can create a Qdrant cluster in that enviroment. Use the same process to [Create a cluster](/documentation/cloud/create-cluster/). Make sure to select your Hybrid Cloud Environment as the target.
|
||||
|
||||

|
||||
|
||||
Note that in the "Kubernetes Configuration" section you can additionally configure:
|
||||
|
||||
* Node selectors for the Qdrant database pods
|
||||
@@ -16,12 +18,20 @@ Note that in the "Kubernetes Configuration" section you can additionally configu
|
||||
|
||||
These settings can also be changed after the cluster is created on the cluster detail page.
|
||||
|
||||

|
||||
|
||||
### Scheduling Configuration
|
||||
|
||||
When creating or editing a cluster, you can configure how the database Pods get scheduled in your Kubernetes cluster. This can be useful to ensure that the Qdrant databases will run on dedicated nodes. You can configure the necessary node selectors and tolerations in the "Kubernetes Configuration" section during cluster creation, or on the cluster detail page.
|
||||
|
||||
### Authentication to your Qdrant clusters
|
||||
|
||||
In Hybrid Cloud the authentication information is provided by Kubernetes secrets.
|
||||
|
||||
You can configure authentication for your Qdrant clusters in the "Configuration" section of the Qdrant Cluster detail page. There you can configure the Kubernetes secret name and key to be used as an API key and/or read-only API key.
|
||||
|
||||

|
||||
|
||||
One way to create a secret is with kubectl:
|
||||
|
||||
```shell
|
||||
@@ -122,7 +132,7 @@ spec:
|
||||
number: 6333
|
||||
```
|
||||
|
||||
Please refer to the Kubernetes, ingress controller and cloud provider documention for more details.
|
||||
Please refer to the Kubernetes, ingress controller and cloud provider documentation for more details.
|
||||
|
||||
If you expose the database like this, you will be able to see this also reflected as an endpoint on the cluster detail page. And will see the Qdrant database dashboard link pointing to it.
|
||||
|
||||
@@ -133,6 +143,8 @@ If you want to configure TLS for accessing your Qdrant database in Hybrid Cloud,
|
||||
* You can offload TLS at the ingress or loadbalancer level.
|
||||
* You can configure TLS directly in the Qdrant database.
|
||||
|
||||
If you want to offload TLS at the ingress or loadbancer level, please refer to their respective documents.
|
||||
|
||||
If you want to configure TLS directly in the Qdrant database, you can reference a secret containing the TLS certificate and key in the "Configuration" section of the Qdrant Cluster detail page.
|
||||
|
||||
To create such a secret, you can use `kubectl`:
|
||||
@@ -155,4 +167,20 @@ metadata:
|
||||
type: kubernetes.io/tls
|
||||
```
|
||||
|
||||
With this command the secret name to enter into the UI would be `qdrant-tls` and the keys would be `tls.crt` and `tls.key`.
|
||||
With this command the secret name to enter into the UI would be `qdrant-tls` and the keys would be `tls.crt` and `tls.key`.
|
||||
|
||||
### Configuring CPU and memory resource reservations
|
||||
|
||||
When creating a Qdrant database cluster, Qdrant Cloud schedules Pods with specific CPU and memory requests and limits to ensure optimal performance. It will use equal requests and limits for stability. Ideally, Kubernetes nodes should match the Pod size, with one database Pod per VM.
|
||||
|
||||
By default, Qdrant Cloud will reserve 20% of available CPU and memory on each Pod. This is done to leave room for the operating system, Kubernetes, and system components. This conservative default may need adjustment depending on node size, whereby smaller nodes might require more, and larger nodes less resources reserved.
|
||||
|
||||
You can modify this reservation in the “Configuration” section of the Qdrant Cluster detail page.
|
||||
|
||||
If you want to check how much resources are availabe on an empty Kubernetes node, you can use the following command:
|
||||
|
||||
```shell
|
||||
kubectl describe node <node-name>
|
||||
```
|
||||
|
||||
This will give you a breakdown of the available resources to Kubernetes and how much is already reserved and used for system Pods.
|
||||
@@ -7,15 +7,19 @@ weight: 1
|
||||
|
||||
The following instruction set will show you how to properly set up a **Qdrant cluster** in your **Hybrid Cloud Environment**.
|
||||
|
||||
You can also watch a video demo on how to set up a Hybrid Cloud Environment:
|
||||
<p align="center"><iframe width="560" height="315" src="https://www.youtube.com/embed/BF02jULGCfo?si=apU1uQOE8AMjq9hD" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></p>
|
||||
|
||||
To learn how Hybrid Cloud works, [read the overview document](/documentation/hybrid-cloud/).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- **Kubernetes cluster:** To create a Hybrid Cloud Environment, you need a [standard compliant](https://www.cncf.io/training/certification/software-conformance/) Kubernetes cluster. You can run this cluster in any cloud, on-premise or edge environment, with distributions that range from AWS EKS to VMWare vSphere. See [Deployment Platforms](/documentation/hybrid-cloud/platform-deployment-options/) for more information.
|
||||
- **Storage:** For storage, you need to set up the Kubernetes cluster with a Container Storage Interface (CSI) driver that provides block storage. For vertical scaling, the CSI driver needs to support volume expansion. For backups and restores, the driver needs to support CSI snapshots and restores.
|
||||
- **Storage:** For storage, you need to set up the Kubernetes cluster with a Container Storage Interface (CSI) driver that provides block storage. For vertical scaling, the CSI driver needs to support volume expansion. The `StorageClass` needs to be created beforehand. For backups and restores, the driver needs to support CSI snapshots and restores. The `VolumeSnapshotClass` needs to be created beforehand. See [Deployment Platforms](/documentation/hybrid-cloud/platform-deployment-options/) for more information.
|
||||
|
||||
<aside role="status">Network storage systems like NFS or object storage systems such as S3 are not supported.</aside>
|
||||
|
||||
- **Kubernetes nodes:** You need enough CPU and memory capacity for the Qdrant database clusters that you create. A small amount of resources is also needed for the Hybrid Cloud control plane components. Qdrant Hybrid Cloud supports x86_64 and ARM64 architectures.
|
||||
- **Permissions:** To install the Qdrant Kubernetes Operator you need to have `cluster-admin` access in your Kubernetes cluster.
|
||||
- **Connection:** The Qdrant Kubernetes Operator in your cluster needs to be able to connect to Qdrant Cloud. It will create an outgoing connection to `cloud.qdrant.io` on port `443`.
|
||||
- **Locations:** By default, the Qdrant Cloud Agent and Operator pulls Helm charts and container images from `registry.cloud.qdrant.io`. The Qdrant database container image is pulled from `docker.io`.
|
||||
@@ -31,13 +35,69 @@ During the onboarding, you will need to deploy the Qdrant Kubernetes Operator an
|
||||
|
||||
You will need to have access to the Kubernetes cluster with `kubectl` and `helm` configured to connect to it. Please refer the documentation of your Kubernetes distribution for more information.
|
||||
|
||||
### Required artifacts
|
||||
## Installation
|
||||
|
||||
1. To set up Hybrid Cloud, open the Qdrant Cloud Console at [cloud.qdrant.io](https://cloud.qdrant.io). On the dashboard, select **Hybrid Cloud**.
|
||||
|
||||
2. Before creating your first Hybrid Cloud Environment, you have to provide billing information and accept the Hybrid Cloud license agreement. The installation wizard will guide you through the process.
|
||||
|
||||
> **Note:** You will only be charged for the Qdrant cluster you create in a Hybrid Cloud Environment, but not for the environment itself.
|
||||
|
||||
3. Now you can specify the following:
|
||||
|
||||
- **Name:** A name for the Hybrid Cloud Environment
|
||||
- **Kubernetes Namespace:** The Kubernetes namespace for the operator and agent. Once you select a namespace, you can't change it.
|
||||
|
||||
You can also configure the StorageClass and VolumeSnapshotClass to use for the Qdrant databases, if you want to deviate from the default settings of your cluster.
|
||||
|
||||

|
||||
|
||||
4. You can then enter the YAML configuration for your Kubernetes operator. Qdrant supports a specific list of configuration options, as described in the [Qdrant Operator configuration](/documentation/hybrid-cloud/operator-configuration/) section.
|
||||
|
||||
5. (Optional) If you have special requirements for any of the following, activate the **Show advanced configuration** option:
|
||||
|
||||
- If you use a proxy to connect from your infrastructure to the Qdrant Cloud API, you can specify the proxy URL, credentials and cetificates.
|
||||
- Container registry URL for Qdrant Operator and Agent images. The default is <https://registry.cloud.qdrant.io/qdrant/>.
|
||||
- Helm chart repository URL for the Qdrant Operator and Agent. The default is <oci://registry.cloud.qdrant.io/qdrant-charts>.
|
||||
- An optional secret with credentials to access your own container registry.
|
||||
- Log level for the operator and agent
|
||||
- Node selectors and tolerations for the operater, agent and monitoring stack
|
||||
|
||||

|
||||
|
||||
6. Once complete, click **Create**.
|
||||
|
||||
> **Note:** All settings but the Kubernetes namespace can be changed later.
|
||||
|
||||
### Generate Installation Command
|
||||
|
||||
After creating your Hybrid Cloud, select **Generate Installation Command** to generate a script that you can run in your Kubernetes cluster which will perform the initial installation of the Kubernetes operator and agent.
|
||||
|
||||

|
||||
|
||||
It will:
|
||||
|
||||
- Create the Kubernetes namespace, if not present.
|
||||
- Set up the necessary secrets with credentials to access the Qdrant container registry and the Qdrant Cloud API.
|
||||
- Sign in to the Helm registry at `registry.cloud.qdrant.io`.
|
||||
- Install the Qdrant cloud agent and Kubernetes operator chart.
|
||||
|
||||
You need this command only for the initial installation. After that, you can update the agent and operator using the Qdrant Cloud Console.
|
||||
|
||||
> **Note:** If you generate the installation command a second time, it will re-generate the included secrets, and you will have to apply the command again to update them.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
### Mirroring images and charts
|
||||
|
||||
#### Required artifacts
|
||||
|
||||
Container images:
|
||||
|
||||
- `docker.io/qdrant/qdrant`
|
||||
- `registry.cloud.qdrant.io/qdrant/qdrant`
|
||||
- `registry.cloud.qdrant.io/qdrant/qdrant-cloud-agent`
|
||||
- `registry.cloud.qdrant.io/qdrant/qdrant-operator`
|
||||
- `registry.cloud.qdrant.io/qdrant/operator`
|
||||
- `registry.cloud.qdrant.io/qdrant/cluster-manager`
|
||||
- `registry.cloud.qdrant.io/qdrant/prometheus`
|
||||
- `registry.cloud.qdrant.io/qdrant/prometheus-config-reloader`
|
||||
@@ -47,30 +107,32 @@ Open Containers Initiative (OCI) Helm charts:
|
||||
|
||||
- `registry.cloud.qdrant.io/qdrant-charts/qdrant-cloud-agent`
|
||||
- `registry.cloud.qdrant.io/qdrant-charts/qdrant-operator`
|
||||
- `registry.cloud.qdrant.io/qdrant-charts/operator`
|
||||
- `registry.cloud.qdrant.io/qdrant-charts/qdrant-cluster-manager`
|
||||
- `registry.cloud.qdrant.io/qdrant-charts/prometheus`
|
||||
|
||||
### Rate limits at `docker.io`
|
||||
To mirror all necessary container images and Helm charts into your own registry, you should use an automatic replication feature that your registry provides, so that you have new image versions available automatically. Alternatively you can manually sync the images with tools like [Skopeo](https://github.com/containers/skopeo). When syncing images manually, make sure that you sync then with all, or with the right CPU architecture.
|
||||
|
||||
By default, the Qdrant database image will be fetched from Docker Hub, which is the main source of truth. Docker Hub has rate limits for anonymous users. If you have larger setups and also fetch other images from their, you may run into these limits. To solve this, you can provide authentication information for Docker Hub.
|
||||
##### Automatic replication
|
||||
|
||||
First, create a secret with your Docker Hub credentials into your `the-qdrant-namespace` namespace:
|
||||
Ensure that you have both the container images in the `/qdrant/` repository, and the helm charts in the `/qdrant-charts/` repository synced. Then go to the advanced section of your Hybrid Cloud Environment and configure your registry locations:
|
||||
|
||||
* Container registry URL: `your-registry.example.com/qdrant` (this will for example result in `your-registry.example.com/qdrant/qdrant-cloud-agent`)
|
||||
* Chart repository URL: `oci://your-registry.example.com/qdrant-charts` (this will for example result in `oci://your-registry.example.com/qdrant-charts/qdrant-cloud-agent`)
|
||||
|
||||
If you registry requires authentication, you have to create your own secrets with authentication information into your `the-qdrant-namespace` namespace.
|
||||
|
||||
Example:
|
||||
|
||||
```shell
|
||||
kubectl create secret docker-registry dockerhub-registry-secret --namespace the-qdrant-namespace --docker-server=https://index.docker.io/v1/ --docker-username=<your-name> --docker-password=<your-pword> --docker-email=<your-email>
|
||||
kubectl --namespace the-qdrant-namespace create secret docker-registry my-creds --docker-server='your-registry.example.com' --docker-username='your-username' --docker-password='your-password'
|
||||
```
|
||||
|
||||
Then, you can reference this secret by adding the following configuration in the operator configuration YAML editor in the advanced section of the Hybrid Cloud Environment:
|
||||
You can then reference they secret in the advanced section of your Hybrid Cloud Environment.
|
||||
|
||||
```yaml
|
||||
qdrant:
|
||||
image:
|
||||
pull_secret: "dockerhub-registry-secret"
|
||||
```
|
||||
##### Manual replication
|
||||
|
||||
### Mirroring images and charts
|
||||
|
||||
To mirror all necessary container images and Helm charts into your own registry, you can either use a replication feature that your registry provides, or you can manually sync the images with [Skopeo](https://github.com/containers/skopeo):
|
||||
This example uses Skopeo.
|
||||
|
||||
You can find your personal credentials for the Qdrant Cloud registry in the onboarding command, or you can fetch them with `kubectl`:
|
||||
|
||||
@@ -94,6 +156,7 @@ To sync all container images:
|
||||
|
||||
```shell
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant/qdrant-operator your-registry.example.com/qdrant/qdrant-operator
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant/operator your-registry.example.com/qdrant/operator
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant/qdrant-cloud-agent your-registry.example.com/qdrant/qdrant-cloud-agent
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant/prometheus your-registry.example.com/qdrant/prometheus
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant/prometheus-config-reloader your-registry.example.com/qdrant/prometheus-config-reloader
|
||||
@@ -107,8 +170,9 @@ To sync all helm charts:
|
||||
|
||||
```shell
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/prometheus your-registry.example.com/qdrant-charts/prometheus
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/operator your-registry.example.com/qdrant-charts/operator
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/qdrant-operator your-registry.example.com/qdrant-charts/qdrant-operator
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/qdrant-operator-crds your-registry.example.com/qdrant-charts/qdrant-operator-crds
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api your-registry.example.com/qdrant-charts/qdrant-kubernetes-api
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/qdrant-cloud-agent your-registry.example.com/qdrant-charts/qdrant-cloud-agent
|
||||
skopeo sync --all --src docker --dest docker registry.cloud.qdrant.io/qdrant-charts/operator your-registry.example.com/qdrant-charts/operator
|
||||
```
|
||||
@@ -118,68 +182,47 @@ With the above configuration, you can add the following values to the advanced s
|
||||
* Container registry URL: `your-registry.example.com/qdrant`
|
||||
* Chart repository URL: `oci://your-registry.example.com/qdrant-charts`
|
||||
|
||||
If you registry requires authentication, you have to create your own secrets with authentication information into your `the-qdrant-namespace` namespace.
|
||||
If your registry requires authentication, you can create and reference the secret the same way as described above.
|
||||
|
||||
You can then reference they secret by
|
||||
### Rate limits at `docker.io`
|
||||
|
||||
## Installation
|
||||
By default, the Qdrant database image will be fetched from Docker Hub, which is the main source of truth. Docker Hub has rate limits for anonymous users. If you have larger setups and also fetch other images from their, you may run into these limits. To solve this, you can provide authentication information for Docker Hub.
|
||||
|
||||
1. To set up Hybrid Cloud, open the Qdrant Cloud Console at [cloud.qdrant.io](https://cloud.qdrant.io). On the dashboard, select **Hybrid Cloud**.
|
||||
First, create a secret with your Docker Hub credentials into your `the-qdrant-namespace` namespace:
|
||||
|
||||
2. Before creating your first Hybrid Cloud Environment, you have to provide billing information and accept the Hybrid Cloud license agreement. The installation wizard will guide you through the process.
|
||||
```shell
|
||||
kubectl create secret docker-registry dockerhub-registry-secret --namespace the-qdrant-namespace --docker-server=https://index.docker.io/v1/ --docker-username=<your-name> --docker-password=<your-pword> --docker-email=<your-email>
|
||||
```
|
||||
|
||||
> **Note:** You will only be charged for the Qdrant cluster you create in a Hybrid Cloud Environment, but not for the environment itself.
|
||||
Then, you can reference this secret by adding the following configuration in the operator configuration YAML editor in the advanced section of the Hybrid Cloud Environment:
|
||||
|
||||
3. Now you can specify the following:
|
||||
```yaml
|
||||
qdrant:
|
||||
image:
|
||||
pull_secret: "dockerhub-registry-secret"
|
||||
```
|
||||
|
||||
- **Name:** A name for the Hybrid Cloud Environment
|
||||
- **Kubernetes Namespace:** The Kubernetes namespace for the operator and agent. Once you select a namespace, you can't change it.
|
||||
## Rotating Secrets
|
||||
|
||||
You can also configure the StorageClass and VolumeSnapshotClass to use for the Qdrant databases, if you want to deviate from the default settings of your cluster.
|
||||
If you need to rotate the secrets to pull container images and charts from the Qdrant registry and to authenticate at the Qdrant Cloud API, you can do so by following these steps:
|
||||
|
||||
4. You can then enter the YAML configuration for your Kubernetes operator. Qdrant supports a specific list of configuration options, as described in the [Qdrant Operator configuration](/documentation/hybrid-cloud/operator-configuration/) section.
|
||||
* Go to the Hybrid Cloud environment list or the detail page of the environment.
|
||||
* In the actions menu, choose "Rotate Secrets"
|
||||
* Confirm the action
|
||||
* You will receive a new installation command that you can run in your Kubernetes cluster to update the secrets.
|
||||
|
||||
5. (Optional) If you have special requirements for any of the following, activate the **Show advanced configuration** option:
|
||||
If you don't run the installation command, the secrets will not be updated and the communication between your Hybrid Cloud Environment and the Qdrant Cloud API will not work.
|
||||
|
||||
- If you use a proxy to connect from your infrastructure to the Qdrant Cloud API, you can specify the proxy URL, credentials and cetificates.
|
||||
- Container registry URL for Qdrant Operator and Agent images. The default is <https://registry.cloud.qdrant.io/qdrant/>.
|
||||
- Helm chart repository URL for the Qdrant Operator and Agent. The default is <oci://registry.cloud.qdrant.io/qdrant-charts>.
|
||||
- Log level for the operator and agent
|
||||
|
||||
6. Once complete, click **Create**.
|
||||
|
||||
> **Note:** All settings but the Kubernetes namespace can be changed later.
|
||||
|
||||
### Generate Installation Command
|
||||
|
||||
After creating your Hybrid Cloud, select **Generate Installation Command** to generate a script that you can run in your Kubernetes cluster which will perform the initial installation of the Kubernetes operator and agent. It will:
|
||||
|
||||
- Create the Kubernetes namespace, if not present
|
||||
- Set up the necessary secrets with credentials to access the Qdrant container registry and the Qdrant Cloud API.
|
||||
- Sign in to the Helm registry at `registry.cloud.qdrant.io`
|
||||
- Install the Qdrant cloud agent and Kubernetes operator chart
|
||||
|
||||
You need this command only for the initial installation. After that, you can update the agent and operator using the Qdrant Cloud Console.
|
||||
|
||||
> **Note:** If you generate the installation command a second time, it will re-generate the included secrets, and you will have to apply the command again to update them.
|
||||

|
||||
|
||||
## Deleting a Hybrid Cloud Environment
|
||||
|
||||
To delete a Hybrid Cloud Environment, first delete all Qdrant database clusters in it. Then you can delete the environment itself.
|
||||
|
||||
To clean up your Kubernetes cluster, after deleting the Hybrid Cloud Environment, you can use the following command:
|
||||
To clean up your Kubernetes cluster, after deleting the Hybrid Cloud Environment, you can use the following script to remove all Qdrant related resources:
|
||||
|
||||
Download the script at https://github.com/qdrant/qdrant-cloud-support-tools/tree/main/hybrid-cloud-cleanup. Then run the following command while being connected to your Kubernetes cluster. The script requires `kubectl` and `helm` to be installed.
|
||||
|
||||
```shell
|
||||
helm -n the-qdrant-namespace delete qdrant-cloud-agent
|
||||
helm -n the-qdrant-namespace delete qdrant-prometheus
|
||||
helm -n the-qdrant-namespace delete qdrant-operator
|
||||
kubectl -n the-qdrant-namespace patch HelmRelease.cd.qdrant.io qdrant-cloud-agent -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl -n the-qdrant-namespace patch HelmRelease.cd.qdrant.io qdrant-prometheus -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl -n the-qdrant-namespace patch HelmRelease.cd.qdrant.io qdrant-operator -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl -n the-qdrant-namespace patch HelmChart.cd.qdrant.io the-qdrant-namespace-qdrant-cloud-agent -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl -n the-qdrant-namespace patch HelmChart.cd.qdrant.io the-qdrant-namespace-qdrant-prometheus -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl -n the-qdrant-namespace patch HelmChart.cd.qdrant.io the-qdrant-namespace-qdrant-operator -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl -n the-qdrant-namespace patch HelmRepository.cd.qdrant.io qdrant-cloud -p '{"metadata":{"finalizers":null}}' --type=merge
|
||||
kubectl delete namespace the-qdrant-namespace
|
||||
kubectl get crd -o name | grep qdrant | xargs -n 1 kubectl delete
|
||||
./hybrid-cloud-cleanup.sh your-qdrant-namespace
|
||||
```
|
||||
|
||||
@@ -33,6 +33,11 @@ At the time of writing, Linode [does not support CSI Volume Snapshots](https://g
|
||||
|
||||
First, consult AWS' managed Kubernetes instructions below. Then, **to set up Qdrant Hybrid Cloud on AWS**, follow our [step-by-step documentation](/documentation/hybrid-cloud/hybrid-cloud-setup/).
|
||||
|
||||
For a good balance between peformance and cost, we recommend:
|
||||
|
||||
* Depending on your cluster resource configuration either general purpose (m6*, m7*, or m8*), memory optimized (r6*, r7*, or r8*) or cpu optimized (c6*, c7*, or c8*) instance types. Qdrant Hybrid Cloud also supports AWS Graviton ARM64 instances.
|
||||
* At least gp3 EBS volumes for storage
|
||||
|
||||
### More on Amazon Elastic Kubernetes Service
|
||||
|
||||
- [Getting Started with Amazon EKS](https://docs.aws.amazon.com/eks/)
|
||||
@@ -131,6 +136,11 @@ First, consult Gcore's managed Kubernetes instructions below. Then, **to set up
|
||||
|
||||
First, consult GCP's managed Kubernetes instructions below. Then, **to set up Qdrant Hybrid Cloud on GCP**, follow our [step-by-step documentation](/documentation/hybrid-cloud/hybrid-cloud-setup/).
|
||||
|
||||
For a good balance between peformance and cost, we recommend:
|
||||
|
||||
* Depending on your cluster resource configuration either general purpose (standard), memory optimized (highmem) or cpu optimized (highcpu) instance types of at least 2nd generation. Qdrant Hybrid Cloud also supports ARM64 instances.
|
||||
* At least pd-balanced disks for storage
|
||||
|
||||
### More on the Google Kubernetes Engine
|
||||
|
||||
- [Getting Started with GKE](https://cloud.google.com/kubernetes-engine/docs/quickstart)
|
||||
@@ -157,6 +167,11 @@ With [Azure Kubernetes Service (AKS)](https://azure.microsoft.com/en-in/products
|
||||
|
||||
First, consult Azure's managed Kubernetes instructions below. Then, **to set up Qdrant Hybrid Cloud on Azure**, follow our [step-by-step documentation](/documentation/hybrid-cloud/hybrid-cloud-setup/).
|
||||
|
||||
For a good balance between peformance and cost, we recommend:
|
||||
|
||||
* Depending on your cluster resource configuration either general purpose (D-family), memory optimized (E-family) or cpu optimized (F-family) instance types. Qdrant Hybrid Cloud also supports Azure Cobalt ARM64 instances.
|
||||
* At least Premium SSD v2 disks for storage
|
||||
|
||||
### More on Azure Kubernetes Service
|
||||
|
||||
- [Getting Started with AKS](https://learn.microsoft.com/en-us/azure/architecture/reference-architectures/containers/aks-start-here)
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user