docs: Clean up integration docs (#1840)
* docs: Clean up integration docs Signed-off-by: Anush008 <anushshetty90@gmail.com> * chore: More cleanup Signed-off-by: Anush008 <anushshetty90@gmail.com> * chore: Deleted image artifacts Signed-off-by: Anush008 <anushshetty90@gmail.com> * chore: Updated redirects Signed-off-by: Anush008 <anushshetty90@gmail.com> --------- Signed-off-by: Anush008 <anushshetty90@gmail.com>
@@ -51,3 +51,33 @@ HUGO_PARAMS_onetrustScriptId = "0196246a-3663-7350-9a45-b65f645d6314"
|
||||
to = "https://qdrant.to/cloud-terms/"
|
||||
status = 301
|
||||
force = true
|
||||
|
||||
[[redirects]]
|
||||
from = "/documentation/platforms/:slug/"
|
||||
to = "/documentation/platforms/"
|
||||
status = 301
|
||||
force = false
|
||||
|
||||
[[redirects]]
|
||||
from = "/documentation/data-management/:slug/"
|
||||
to = "/documentation/data-management/"
|
||||
status = 301
|
||||
force = false
|
||||
|
||||
[[redirects]]
|
||||
from = "/documentation/frameworks/:slug/"
|
||||
to = "/documentation/frameworks/"
|
||||
status = 301
|
||||
force = false
|
||||
|
||||
[[redirects]]
|
||||
from = "/documentation/embeddings/:slug/"
|
||||
to = "/documentation/embeddings/"
|
||||
status = 301
|
||||
force = false
|
||||
|
||||
[[redirects]]
|
||||
from = "/documentation/observability/:slug/"
|
||||
to = "/documentation/observability/"
|
||||
status = 301
|
||||
force = false
|
||||
@@ -65,8 +65,6 @@ aliases: # There is no need to add aliases for future new tags and categories!
|
||||
- /tags/embedding
|
||||
- /tags/corporate-news
|
||||
- /tags/nvidia
|
||||
- /tags/docarray
|
||||
- /tags/jina-integration
|
||||
- /categories
|
||||
- /categories/news
|
||||
- /categories/vector-search
|
||||
|
||||
@@ -1,28 +0,0 @@
|
||||
---
|
||||
draft: false
|
||||
preview_image: /blog/from_cms/docarray.png
|
||||
sitemapExclude: true
|
||||
title: "Qdrant and Jina integration: storage backend support for DocArray"
|
||||
slug: qdrant-and-jina-integration
|
||||
short_description: "One more way to use Qdrant: Jina's DocArray is now
|
||||
supporting Qdrant as a storage backend."
|
||||
description: We are happy to announce that Jina.AI integrates Qdrant engine as a
|
||||
storage backend to their DocArray solution.
|
||||
date: 2022-03-15T15:00:00+03:00
|
||||
author: Alyona Kavyerina
|
||||
featured: false
|
||||
author_link: https://medium.com/@alyona.kavyerina
|
||||
tags:
|
||||
- jina integration
|
||||
- docarray
|
||||
categories:
|
||||
- News
|
||||
---
|
||||
We are happy to announce that [Jina.AI](https://jina.ai/) integrates Qdrant engine as a storage backend to their [DocArray](https://docarray.jina.ai/) solution.
|
||||
|
||||
Now you can experience the convenience of Pythonic API and Rust performance in a single workflow.
|
||||
|
||||
DocArray library defines a structure for the unstructured data and simplifies processing a collection of documents,
|
||||
including audio, video, text, and other data types. Qdrant engine empowers scaling of its vector search and storage.
|
||||
|
||||
Read more about the integration by this [link](/documentation/install/#docarray)
|
||||
@@ -16,8 +16,5 @@ partition: build
|
||||
| [Confluent](/documentation/data-management/confluent/) | Fully-managed data streaming platform with a cloud-native Apache Kafka engine. |
|
||||
| [DLT](/documentation/data-management/dlt/) | Python library to simplify data loading processes between several sources and destinations. |
|
||||
| [Fluvio](/documentation/data-management/fluvio/) | Rust-based platform for high speed, real-time data processing. |
|
||||
| [Fondant](/documentation/data-management/fondant/) | Framework for developing datasets, sharing reusable operations and data processing trees. |
|
||||
| [MindsDB](/documentation/data-management/mindsdb/) | Platform to deploy, serve, and fine-tune models with numerous data source integrations. |
|
||||
| [NiFi](/documentation/data-management/nifi/) | Data ingestion platform to manage data transfer between different sources and destination systems. |
|
||||
| [Spark](/documentation/data-management/spark/) | A unified analytics engine for large-scale data processing. |
|
||||
| [Unstructured](/documentation/data-management/unstructured/) | Python library with components for ingesting and pre-processing data from numerous sources. |
|
||||
|
||||
@@ -1,81 +0,0 @@
|
||||
---
|
||||
title: Fondant
|
||||
aliases: [ ../integrations/fondant/, ../frameworks/fondant/ ]
|
||||
---
|
||||
|
||||
# Fondant
|
||||
|
||||
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed
|
||||
up large-scale data processing by making containerized components reusable across pipelines and
|
||||
execution environments. Benefit from built-in features such as autoscaling, data lineage, and
|
||||
pipeline caching, and deploy to (managed) platforms such as Vertex AI, Sagemaker, and Kubeflow
|
||||
Pipelines.
|
||||
|
||||
Fondant comes with a library of reusable components that you can leverage to compose your own
|
||||
pipeline, including a Qdrant component for writing embeddings to Qdrant.
|
||||
|
||||
## Usage
|
||||
|
||||
<aside role="status">
|
||||
A Qdrant collection has to be <a href="/documentation/concepts/collections/">created in advance</a>
|
||||
</aside>
|
||||
|
||||
**A data load pipeline for RAG using Qdrant**.
|
||||
|
||||
A simple ingestion pipeline could look like the following:
|
||||
|
||||
```python
|
||||
import pyarrow as pa
|
||||
from fondant.pipeline import Pipeline
|
||||
|
||||
indexing_pipeline = Pipeline(
|
||||
name="ingestion-pipeline",
|
||||
description="Pipeline to prepare and process data for building a RAG solution",
|
||||
base_path="./fondant-artifacts",
|
||||
)
|
||||
|
||||
# An custom implemenation of a read component.
|
||||
text = indexing_pipeline.read(
|
||||
"path/to/data-source-component",
|
||||
arguments={
|
||||
# your custom arguments
|
||||
}
|
||||
)
|
||||
|
||||
chunks = text.apply(
|
||||
"chunk_text",
|
||||
arguments={
|
||||
"chunk_size": 512,
|
||||
"chunk_overlap": 32,
|
||||
},
|
||||
)
|
||||
|
||||
embeddings = chunks.apply(
|
||||
"embed_text",
|
||||
arguments={
|
||||
"model_provider": "huggingface",
|
||||
"model": "all-MiniLM-L6-v2",
|
||||
},
|
||||
)
|
||||
|
||||
embeddings.write(
|
||||
"index_qdrant",
|
||||
arguments={
|
||||
"url": "http:localhost:6333",
|
||||
"collection_name": "some-collection-name",
|
||||
},
|
||||
cache=False,
|
||||
)
|
||||
```
|
||||
|
||||
Once you have a pipeline, you can easily run it using the built-in CLI. Fondant allows
|
||||
you to run the pipeline in production across different clouds.
|
||||
|
||||
The first component is a custom read module that needs to be implemented and cannot be used off the
|
||||
shelf. A detailed tutorial on how to rebuild this
|
||||
pipeline [is provided on GitHub](https://github.com/ml6team/fondant-usecase-RAG/tree/main).
|
||||
|
||||
## Next steps
|
||||
|
||||
More information about creating your own pipelines and components can be found in the [Fondant
|
||||
documentation](https://fondant.ai/en/stable/).
|
||||
@@ -1,97 +0,0 @@
|
||||
---
|
||||
title: MindsDB
|
||||
aliases: [ ../integrations/mindsdb/, ../frameworks/mindsdb/ ]
|
||||
---
|
||||
|
||||
# MindsDB
|
||||
|
||||
[MindsDB](https://mindsdb.com) is an AI automation platform for building AI/ML powered features and applications. It works by connecting any source of data with any AI/ML model or framework and automating how real-time data flows between them.
|
||||
|
||||
With the MindsDB-Qdrant integration, you can now select Qdrant as a database to load into and retrieve from with semantic search and filtering.
|
||||
|
||||
**MindsDB allows you to easily**:
|
||||
|
||||
- Connect to any store of data or end-user application.
|
||||
- Pass data to an AI model from any store of data or end-user application.
|
||||
- Plug the output of an AI model into any store of data or end-user application.
|
||||
- Fully automate these workflows to build AI-powered features and applications
|
||||
|
||||
## Usage
|
||||
|
||||
To get started with Qdrant and MindsDB, the following syntax can be used.
|
||||
|
||||
```sql
|
||||
CREATE DATABASE qdrant_test
|
||||
WITH ENGINE = "qdrant",
|
||||
PARAMETERS = {
|
||||
"location": ":memory:",
|
||||
"collection_config": {
|
||||
"size": 386,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The available arguments for instantiating Qdrant can be found [here](https://github.com/mindsdb/mindsdb/blob/23a509cb26bacae9cc22475497b8644e3f3e23c3/mindsdb/integrations/handlers/qdrant_handler/qdrant_handler.py#L408-L468).
|
||||
|
||||
## Creating a new table
|
||||
|
||||
- Qdrant options for creating a collection can be specified as `collection_config` in the `CREATE DATABASE` parameters.
|
||||
- By default, UUIDs are set as collection IDs. You can provide your own IDs under the `id` column.
|
||||
|
||||
```sql
|
||||
CREATE TABLE qdrant_test.test_table (
|
||||
SELECT embeddings,'{"source": "bbc"}' as metadata FROM mysql_demo_db.test_embeddings
|
||||
);
|
||||
```
|
||||
|
||||
## Querying the database
|
||||
|
||||
#### Perform a full retrieval using the following syntax.
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
```
|
||||
|
||||
By default, the `LIMIT` is set to 10 and the `OFFSET` is set to 0.
|
||||
|
||||
#### Perform a similarity search using your embeddings
|
||||
|
||||
<aside role="status">Qdrant supports <a href="/documentation/concepts/indexing/#payload-index">payload indexing</a> that vastly improves retrieval efficiency with filters and is highly recommended. Please note that this feature currently cannot be configured via MindsDB and must be set up separately if needed.</aside>
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
WHERE search_vector = (select embeddings from mysql_demo_db.test_embeddings limit 1)
|
||||
```
|
||||
|
||||
#### Perform a search using filters
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
WHERE `metadata.source` = 'bbc';
|
||||
```
|
||||
|
||||
#### Delete entries using IDs
|
||||
|
||||
```sql
|
||||
DELETE FROM qtest.test_table_6
|
||||
WHERE id = 2
|
||||
```
|
||||
|
||||
#### Delete entries using filters
|
||||
|
||||
```sql
|
||||
DELETE * FROM qdrant_test.test_table
|
||||
WHERE `metadata.source` = 'bbc';
|
||||
```
|
||||
|
||||
#### Drop a table
|
||||
|
||||
```sql
|
||||
DROP TABLE qdrant_test.test_table;
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- You can find more information pertaining to MindsDB and its datasources [here](https://docs.mindsdb.com/).
|
||||
- [Source Code](https://github.com/mindsdb/mindsdb/tree/main/mindsdb/integrations/handlers/qdrant_handler)
|
||||
@@ -1,33 +0,0 @@
|
||||
---
|
||||
title: Apache NiFi
|
||||
aliases: [ ../frameworks/nifi/ ]
|
||||
---
|
||||
|
||||
# Apache NiFi
|
||||
|
||||
[NiFi](https://nifi.apache.org/) is a real-time data ingestion platform, which can transfer and manage data transfer between numerous sources and destination systems. It supports many protocols and offers a web-based user interface for developing and monitoring data flows.
|
||||
|
||||
NiFi supports ingesting and querying data in Qdrant via its processor modules.
|
||||
|
||||
## Configuration
|
||||
|
||||

|
||||
|
||||
You can configure Qdrant NiFi processors with your Qdrant credentials, query/upload configurations. The processors offer 2 built-in embedding providers to encode data into vector embeddings - HuggingFace, OpenAI.
|
||||
|
||||
## Put Qdrant
|
||||
|
||||

|
||||
|
||||
The `Put Qdrant` processor can ingest NiFi [FlowFile](https://nifi.apache.org/docs/nifi-docs/html/nifi-in-depth.html#intro) data into a Qdrant collection.
|
||||
|
||||
## Query Qdrant
|
||||
|
||||

|
||||
|
||||
The `Query Qdrant` processor can perform a similarity search across a Qdrant collection and return a [FlowFile](https://nifi.apache.org/docs/nifi-docs/html/nifi-in-depth.html#intro) result.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [NiFi Documentation](https://nifi.apache.org/documentation/v2/).
|
||||
- [Source Code](https://github.com/apache/nifi-python-extensions)
|
||||
@@ -14,7 +14,7 @@ Qdrant can be used as an ingestion destination in Unstructured.
|
||||
Install Unstructured with the `qdrant` extra.
|
||||
|
||||
```bash
|
||||
pip install "unstructured[qdrant]"
|
||||
pip install "unstructured-ingest[qdrant]"
|
||||
```
|
||||
|
||||
## Usage
|
||||
@@ -25,21 +25,21 @@ Depending on the use case you can prefer the command line or using it within you
|
||||
### CLI
|
||||
|
||||
```bash
|
||||
EMBEDDING_PROVIDER=${EMBEDDING_PROVIDER:-"langchain-huggingface"}
|
||||
|
||||
unstructured-ingest \
|
||||
local \
|
||||
--input-path example-docs/book-war-and-peace-1225p.txt \
|
||||
--output-dir local-output-to-qdrant \
|
||||
--strategy fast \
|
||||
--chunk-elements \
|
||||
--embedding-provider "$EMBEDDING_PROVIDER" \
|
||||
--num-processes 2 \
|
||||
--verbose \
|
||||
qdrant \
|
||||
--collection-name "test" \
|
||||
--url "http://localhost:6333" \
|
||||
--batch-size 80
|
||||
--input-path $LOCAL_FILE_INPUT_DIR \
|
||||
--chunking-strategy by_title \
|
||||
--embedding-provider huggingface \
|
||||
--partition-by-api \
|
||||
--api-key $UNSTRUCTURED_API_KEY \
|
||||
--partition-endpoint $UNSTRUCTURED_API_URL \
|
||||
--additional-partition-args="{\"split_pdf_page\":\"true\", \"split_pdf_allow_failed\":\"true\", \"split_pdf_concurrency_level\": 15}" \
|
||||
qdrant-cloud \
|
||||
--url $QDRANT_URL \
|
||||
--api-key $QDRANT_API_KEY \
|
||||
--collection-name $QDRANT_COLLECTION \
|
||||
--batch-size 50 \
|
||||
--num-processes 1
|
||||
```
|
||||
|
||||
For a full list of the options the CLI accepts, run `unstructured-ingest <upstream connector> qdrant --help`
|
||||
@@ -47,54 +47,63 @@ For a full list of the options the CLI accepts, run `unstructured-ingest <upstre
|
||||
### Programmatic usage
|
||||
|
||||
```python
|
||||
from unstructured.ingest.connector.local import SimpleLocalConfig
|
||||
from unstructured.ingest.connector.qdrant import (
|
||||
QdrantWriteConfig,
|
||||
SimpleQdrantConfig,
|
||||
)
|
||||
from unstructured.ingest.interfaces import (
|
||||
ChunkingConfig,
|
||||
EmbeddingConfig,
|
||||
PartitionConfig,
|
||||
ProcessorConfig,
|
||||
ReadConfig,
|
||||
)
|
||||
from unstructured.ingest.runner import LocalRunner
|
||||
from unstructured.ingest.runner.writers.base_writer import Writer
|
||||
from unstructured.ingest.runner.writers.qdrant import QdrantWriter
|
||||
import os
|
||||
|
||||
def get_writer() -> Writer:
|
||||
return QdrantWriter(
|
||||
connector_config=SimpleQdrantConfig(
|
||||
url="http://localhost:6333",
|
||||
collection_name="test",
|
||||
),
|
||||
write_config=QdrantWriteConfig(batch_size=80),
|
||||
)
|
||||
from unstructured_ingest.pipeline.pipeline import Pipeline
|
||||
from unstructured_ingest.interfaces import ProcessorConfig
|
||||
|
||||
from unstructured_ingest.processes.connectors.local import (
|
||||
LocalIndexerConfig,
|
||||
LocalDownloaderConfig,
|
||||
LocalConnectionConfig
|
||||
)
|
||||
from unstructured_ingest.processes.partitioner import PartitionerConfig
|
||||
from unstructured_ingest.processes.chunker import ChunkerConfig
|
||||
from unstructured_ingest.processes.embedder import EmbedderConfig
|
||||
|
||||
from unstructured_ingest.processes.connectors.qdrant.cloud import (
|
||||
CloudQdrantConnectionConfig,
|
||||
CloudQdrantAccessConfig,
|
||||
CloudQdrantUploadStagerConfig,
|
||||
CloudQdrantUploaderConfig
|
||||
)
|
||||
|
||||
if __name__ == "__main__":
|
||||
writer = get_writer()
|
||||
runner = LocalRunner(
|
||||
processor_config=ProcessorConfig(
|
||||
verbose=True,
|
||||
output_dir="local-output-to-qdrant",
|
||||
num_processes=2,
|
||||
Pipeline.from_configs(
|
||||
context=ProcessorConfig(),
|
||||
indexer_config=LocalIndexerConfig(input_path=os.getenv("LOCAL_FILE_INPUT_DIR")),
|
||||
downloader_config=LocalDownloaderConfig(),
|
||||
source_connection_config=LocalConnectionConfig(),
|
||||
partitioner_config=PartitionerConfig(
|
||||
partition_by_api=True,
|
||||
api_key=os.getenv("UNSTRUCTURED_API_KEY"),
|
||||
partition_endpoint=os.getenv("UNSTRUCTURED_API_URL"),
|
||||
additional_partition_args={
|
||||
"split_pdf_page": True,
|
||||
"split_pdf_allow_failed": True,
|
||||
"split_pdf_concurrency_level": 15
|
||||
}
|
||||
),
|
||||
connector_config=SimpleLocalConfig(
|
||||
input_path="example-docs/book-war-and-peace-1225p.txt",
|
||||
chunker_config=ChunkerConfig(chunking_strategy="by_title"),
|
||||
embedder_config=EmbedderConfig(embedding_provider="huggingface"),
|
||||
|
||||
destination_connection_config=CloudQdrantConnectionConfig(
|
||||
access_config=CloudQdrantAccessConfig(
|
||||
api_key=os.getenv("QDRANT_API_KEY")
|
||||
),
|
||||
read_config=ReadConfig(),
|
||||
partition_config=PartitionConfig(),
|
||||
chunking_config=ChunkingConfig(chunk_elements=True),
|
||||
embedding_config=EmbeddingConfig(provider="langchain-huggingface"),
|
||||
writer=writer,
|
||||
writer_kwargs={},
|
||||
url=os.getenv("QDRANT_URL")
|
||||
),
|
||||
stager_config=CloudQdrantUploadStagerConfig(),
|
||||
uploader_config=CloudQdrantUploaderConfig(
|
||||
collection_name=os.getenv("QDRANT_COLLECTION"),
|
||||
batch_size=50,
|
||||
num_processes=1
|
||||
)
|
||||
runner.run()
|
||||
).run()
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- Unstructured API [reference](https://unstructured-io.github.io/unstructured/api.html).
|
||||
- Qdrant ingestion destination [reference](https://unstructured-io.github.io/unstructured/ingest/destination_connectors/qdrant.html).
|
||||
- [Source Code](https://github.com/Unstructured-IO/unstructured-ingest/blob/main/unstructured_ingest/connector/qdrant.py)
|
||||
- Qdrant ingestion destination [reference](https://docs.unstructured.io/ui/destinations/qdrant).
|
||||
- [Source Code](https://github.com/Unstructured-IO/unstructured-ingest/tree/main/unstructured_ingest/processes/connectors/qdrant)
|
||||
|
||||
@@ -11,14 +11,11 @@ aliases: ["/documentation/frameworks/memgpt/"]
|
||||
| ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
|
||||
| [AutoGen](/documentation/frameworks/autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
|
||||
| [Camel](/documentation/frameworks/camel/) | Framework to build and use LLM-based agents for real-world task solving |
|
||||
| [Canopy](/documentation/frameworks/canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
|
||||
| [Cheshire Cat](/documentation/frameworks/cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
|
||||
| [CrewAI](/documentation/frameworks/crewai/) | CrewAI is a framework to build automated workflows using multiple AI agents that perform complex tasks. |
|
||||
| [Dagster](/documentation/frameworks/dagster/) | Python framework for data orchestration with integrated lineage, observability. |
|
||||
| [DeepEval](/documentation/frameworks/deepeval/) | Python framework for testing large language model systems. |
|
||||
| [DocArray](/documentation/frameworks/docarray/) | Python library for managing data in multi-modal AI applications. |
|
||||
| [DSPy](/documentation/frameworks/dspy/) | Framework for algorithmically optimizing LM prompts and weights. |
|
||||
| [dsRAG](/documentation/frameworks/dsrag/) | High-performance Python retrieval engine for unstructured data. |
|
||||
| [Dynamiq](/documentation/frameworks/dynamiq/) | Dynamiq is all-in-one Gen AI framework, designed to streamline the development of AI-powered applications. |
|
||||
| [Feast](/documentation/frameworks/feast/) | Open-source feature store to operate production ML systems at scale as a set of features. |
|
||||
| [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. |
|
||||
@@ -27,7 +24,6 @@ aliases: ["/documentation/frameworks/memgpt/"]
|
||||
| [HoneyHive](/documentation/frameworks/honeyhive/) | AI observability and evaluation platform that provides tracing and monitoring tools for GenAI pipelines. |
|
||||
| [Lakechain](/documentation/frameworks/lakechain/) | Python framework for deploying document processing pipelines on AWS using infrastructure-as-code. |
|
||||
| [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [Langchain4j](/documentation/frameworks/langchain4j/) | Java framework for building context-aware, reasoning applications using LLMs. |
|
||||
| [LangGraph](/documentation/frameworks/langgraph/) | Python, Javascript libraries for building stateful, multi-actor applications. |
|
||||
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
|
||||
@@ -36,15 +32,10 @@ aliases: ["/documentation/frameworks/memgpt/"]
|
||||
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
|
||||
| [Neo4j GraphRAG](/documentation/frameworks/neo4j-graphrag/) | Package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. |
|
||||
| [NLWeb](/documentation/frameworks/nlweb/) | A framework to turn websites into chat-ready data using schema.org and associated data formats. |
|
||||
| [OpenAI Agents](/documentation/frameworks/openai-agents/) | Python framework for managing multiple AI agents that can work together. |
|
||||
| [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language |
|
||||
| [Ragbits](/documentation/frameworks/ragbits/) | Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. |
|
||||
| [Rig-rs](/documentation/frameworks/rig-rs/) | Rust library for building scalable, modular, and ergonomic LLM-powered applications. |
|
||||
| [Semantic Router](/documentation/frameworks/semantic-router/) | Python library to build a decision-making layer for AI applications using vector search. |
|
||||
| [SmolAgents](/documentation/frameworks/smolagents/) | Barebones library for agents. Agents write python code to call tools and orchestrate other agent. |
|
||||
| [Solon](/documentation/frameworks/solon/) | A lightweight, high-performance Java enterprise framework |
|
||||
| [Spring AI](/documentation/frameworks/spring-ai/) | Java AI framework for building with Spring design principles such as portability and modular design. |
|
||||
| [Superduper](/documentation/frameworks/superduper/) | Framework for building flexible, compositional AI apps which may be applied directly to databases. |
|
||||
| [Sycamore](/documentation/frameworks/sycamore/) | Document processing engine for ETL, RAG, LLM-based applications, and analytics on unstructured data. |
|
||||
| [Testcontainers](/documentation/frameworks/testcontainers/) | Framework for providing throwaway, lightweight instances of systems for testing |
|
||||
| [txtai](/documentation/frameworks/txtai/) | Python library for semantic search, LLM orchestration and language model workflows. |
|
||||
|
||||
@@ -1,90 +0,0 @@
|
||||
---
|
||||
title: Pinecone Canopy
|
||||
---
|
||||
|
||||
# Pinecone Canopy
|
||||
|
||||
[Canopy](https://github.com/pinecone-io/canopy) is an open-source framework and context engine to build chat assistants at scale.
|
||||
|
||||
Qdrant is supported as a knowledge base within Canopy for context retrieval and augmented generation.
|
||||
|
||||
## Usage
|
||||
|
||||
Install the SDK with the Qdrant extra as described in the [Canopy README](https://github.com/pinecone-io/canopy?tab=readme-ov-file#extras).
|
||||
|
||||
```bash
|
||||
pip install canopy-sdk[qdrant]
|
||||
```
|
||||
|
||||
### Creating a knowledge base
|
||||
|
||||
```python
|
||||
from canopy.knowledge_base import QdrantKnowledgeBase
|
||||
|
||||
kb = QdrantKnowledgeBase(collection_name="<YOUR_COLLECTION_NAME>")
|
||||
```
|
||||
|
||||
<aside role="status">The constructor accepts additional <a href="https://github.com/qdrant/qdrant-client/blob/eda201a1dbf1bbc67415f8437a5619f6f83e8ac6/qdrant_client/qdrant_client.py#L36-L61">options</a> to customize your connection to Qdrant.</aside>
|
||||
|
||||
To create a new Qdrant collection and connect it to the knowledge base, use the `create_canopy_collection` method:
|
||||
|
||||
```python
|
||||
kb.create_canopy_collection()
|
||||
```
|
||||
|
||||
You can always verify the connection to the collection with the `verify_index_connection` method:
|
||||
|
||||
```python
|
||||
kb.verify_index_connection()
|
||||
```
|
||||
|
||||
Learn more about customizing the knowledge base and its inner components [in the Canopy library](https://github.com/pinecone-io/canopy/blob/main/docs/library.md#understanding-knowledgebase-workings).
|
||||
|
||||
### Adding data to the knowledge base
|
||||
|
||||
To insert data into the knowledge base, you can create a list of documents and use the `upsert` method:
|
||||
|
||||
```python
|
||||
from canopy.models.data_models import Document
|
||||
|
||||
documents = [
|
||||
Document(
|
||||
id="1",
|
||||
text="U2 are an Irish rock band from Dublin, formed in 1976.",
|
||||
source="https://en.wikipedia.org/wiki/U2",
|
||||
),
|
||||
Document(
|
||||
id="2",
|
||||
text="Arctic Monkeys are an English rock band formed in Sheffield in 2002.",
|
||||
source="https://en.wikipedia.org/wiki/Arctic_Monkeys",
|
||||
metadata={"my-key": "my-value"},
|
||||
),
|
||||
]
|
||||
|
||||
kb.upsert(documents)
|
||||
```
|
||||
|
||||
### Querying the knowledge base
|
||||
|
||||
You can query the knowledge base with the `query` method to find the most similar documents to a given text:
|
||||
|
||||
```python
|
||||
from canopy.models.data_models import Query
|
||||
|
||||
kb.query(
|
||||
[
|
||||
Query(text="Arctic Monkeys music genre"),
|
||||
Query(
|
||||
text="U2 music genre",
|
||||
top_k=10,
|
||||
metadata_filter={"key": "my-key", "match": {"value": "my-value"}},
|
||||
),
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Introduction to Canopy](https://www.pinecone.io/blog/canopy-rag-framework/)
|
||||
- [Canopy library reference](https://github.com/pinecone-io/canopy/blob/main/docs/library.md)
|
||||
- [Source Code](https://github.com/pinecone-io/canopy/tree/main/src/canopy/knowledge_base/qdrant)
|
||||
@@ -1,22 +0,0 @@
|
||||
---
|
||||
title: DocArray
|
||||
aliases: [ ../integrations/docarray/ ]
|
||||
---
|
||||
|
||||
# DocArray
|
||||
|
||||
You can use Qdrant natively in DocArray, where Qdrant serves as a high-performance document store to enable scalable vector search.
|
||||
|
||||
DocArray is a library from Jina AI for nested, unstructured data in transit, including text, image, audio, video, 3D mesh, etc.
|
||||
It allows deep-learning engineers to efficiently process, embed, search, recommend, store, and transfer the data with a Pythonic API.
|
||||
|
||||
To install DocArray with Qdrant support, please do
|
||||
|
||||
```bash
|
||||
pip install "docarray[qdrant]"
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [DocArray documentations](https://docarray.jina.ai/advanced/document-store/qdrant/).
|
||||
- [Source Code](https://github.com/docarray/docarray/blob/main/docarray/index/backends/qdrant.py)
|
||||
@@ -1,50 +0,0 @@
|
||||
---
|
||||
title: dsRAG
|
||||
---
|
||||
|
||||
# dsRAG
|
||||
|
||||
[dsRAG](https://github.com/D-Star-AI/dsRAG) is a retrieval engine for unstructured data. It is especially good at handling challenging queries over dense text, like financial reports, legal documents, and academic papers. dsRAG achieves substantially higher accuracy than vanilla RAG baselines on complex open-book question answering tasks
|
||||
|
||||
You can use the Qdrant connector in dsRAG to add and semantically retrieve documents from your collections.
|
||||
|
||||
## Usage Example
|
||||
|
||||
```python
|
||||
from dsrag.database.vector import QdrantVectorDB
|
||||
import numpy as np
|
||||
from qdrant_clien import models
|
||||
|
||||
db = QdrantVectorDB(kb_id=self.kb_id, url="http://localhost:6334", prefer_grpc=True)
|
||||
vectors = [np.array([1, 0]), np.array([0, 1])]
|
||||
|
||||
# You can use any document loaders available with dsRAG
|
||||
# We'll use literals for demonstration
|
||||
documents = [
|
||||
{
|
||||
"doc_id": "1",
|
||||
"chunk_index": 0,
|
||||
"chunk_header": "Header1",
|
||||
"chunk_text": "Text1",
|
||||
},
|
||||
{
|
||||
"doc_id": "2",
|
||||
"chunk_index": 1,
|
||||
"chunk_header": "Header2",
|
||||
"chunk_text": "Text2",
|
||||
},
|
||||
]
|
||||
|
||||
db.add_vectors(vectors, documents)
|
||||
|
||||
metadata_filter = models.Filter(
|
||||
must=[models.FieldCondition(key="doc_id", match=models.MatchValue(value="1"))]
|
||||
)
|
||||
|
||||
db.search(query_vector, top_k=4, metadata_filter=metadata_filter)
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [dsRAG Source](https://github.com/D-Star-AI/dsRAG).
|
||||
- [dsRAG Examples](https://github.com/D-Star-AI/dsRAG/tree/main/examples)
|
||||
@@ -59,4 +59,4 @@ feature_values = feature_store.retrieve_online_documents(
|
||||
## 📚 Further Reading
|
||||
|
||||
- [Feast Documentation](http://docs.feast.dev/)
|
||||
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/)
|
||||
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/qdrant_online_store)
|
||||
|
||||
@@ -24,26 +24,23 @@ To use this plugin, specify it when you call `configureGenkit()`:
|
||||
|
||||
```js
|
||||
import { qdrant } from 'genkitx-qdrant';
|
||||
import { textEmbeddingGecko } from '@genkit-ai/vertexai';
|
||||
|
||||
export default configureGenkit({
|
||||
const ai = genkit({
|
||||
plugins: [
|
||||
qdrant([
|
||||
{
|
||||
embedder: googleAI.embedder('text-embedding-004'),
|
||||
collectionName: 'collectionName',
|
||||
clientParams: {
|
||||
host: 'localhost',
|
||||
port: 6333,
|
||||
},
|
||||
collectionName: 'some-collection',
|
||||
embedder: textEmbeddingGecko,
|
||||
},
|
||||
url: 'http://localhost:6333',
|
||||
}
|
||||
}
|
||||
]),
|
||||
],
|
||||
// ...
|
||||
});
|
||||
```
|
||||
|
||||
You'll need to specify a collection name, the embedding model you want to use and the Qdrant client parameters. In
|
||||
You'll need to specify a collection name, the embedding model you want to use and the Qdrant client parameters. In
|
||||
addition, there are a few optional parameters:
|
||||
|
||||
- `embedderOptions`: Additional options to pass options to the embedder:
|
||||
@@ -64,7 +61,13 @@ addition, there are a few optional parameters:
|
||||
metadataPayloadKey: 'metadata';
|
||||
```
|
||||
|
||||
- `collectionCreateOptions`: [Additional options](/documentation/concepts/collections/#create-a-collection/) when creating the Qdrant collection.
|
||||
- `dataTypePayloadKey`: Name of the payload filed with the document datatype. Defaults to "_content_type".
|
||||
|
||||
```js
|
||||
dataTypePayloadKey: '_datatype';
|
||||
```
|
||||
|
||||
- `collectionCreateOptions`: [Additional options](<(https://qdrant.tech/documentation/concepts/collections/#create-a-collection)>) when creating the Qdrant collection.
|
||||
|
||||
## Usage
|
||||
|
||||
@@ -72,36 +75,25 @@ Import retriever and indexer references like so:
|
||||
|
||||
```js
|
||||
import { qdrantIndexerRef, qdrantRetrieverRef } from 'genkitx-qdrant';
|
||||
import { Document, index, retrieve } from '@genkit-ai/ai/retriever';
|
||||
```
|
||||
|
||||
Then, pass the references to `retrieve()` and `index()`:
|
||||
Then, pass their references to `retrieve()` and `index()`:
|
||||
|
||||
```js
|
||||
// To specify an indexer:
|
||||
export const qdrantIndexer = qdrantIndexerRef({
|
||||
collectionName: 'some-collection',
|
||||
displayName: 'Some Collection indexer',
|
||||
});
|
||||
|
||||
await index({ indexer: qdrantIndexer, documents });
|
||||
// To export an indexer reference:
|
||||
export const qdrantIndexer = qdrantIndexerRef('collectionName', 'displayName');
|
||||
```
|
||||
|
||||
```js
|
||||
// To specify a retriever:
|
||||
export const qdrantRetriever = qdrantRetrieverRef({
|
||||
collectionName: 'some-collection',
|
||||
displayName: 'Some Collection Retriever',
|
||||
});
|
||||
|
||||
let docs = await retrieve({ retriever: qdrantRetriever, query });
|
||||
// To export a retriever reference:
|
||||
export const qdrantRetriever = qdrantRetrieverRef('collectionName', 'displayName');
|
||||
```
|
||||
|
||||
You can refer to [Retrieval-augmented generation](https://firebase.google.com/docs/genkit/rag) for a general
|
||||
You can refer to [Retrieval-augmented generation](https://genkit.dev/docs/rag/) for a general
|
||||
discussion on indexers and retrievers.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Introduction to Genkit](https://firebase.google.com/docs/genkit)
|
||||
- [Genkit Documentation](https://firebase.google.com/docs/genkit/get-started)
|
||||
- [Introduction to Genkit](https://genkit.dev/)
|
||||
- [Genkit Documentation](https://genkit.dev/docs/get-started/)
|
||||
- [Source Code](https://github.com/qdrant/qdrant-genkit)
|
||||
|
||||
@@ -1,71 +0,0 @@
|
||||
---
|
||||
title: Langchain Go
|
||||
---
|
||||
|
||||
# Langchain Go
|
||||
|
||||
[Langchain Go](https://tmc.github.io/langchaingo/docs/) is a framework for developing data-aware applications powered by language models in Go.
|
||||
|
||||
You can use Qdrant as a vector store in Langchain Go.
|
||||
|
||||
## Setup
|
||||
|
||||
Install the `langchain-go` project dependency
|
||||
|
||||
```bash
|
||||
go get -u github.com/tmc/langchaingo
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Before you use the following code sample, customize the following values for your configuration:
|
||||
|
||||
- `YOUR_QDRANT_REST_URL`: If you've set up Qdrant using the [Quick Start](/documentation/quick-start/) guide,
|
||||
set this value to `http://localhost:6333`.
|
||||
- `YOUR_COLLECTION_NAME`: Use our [Collections](/documentation/concepts/collections/) guide to create or
|
||||
list collections.
|
||||
|
||||
```go
|
||||
package main
|
||||
|
||||
import (
|
||||
"log"
|
||||
"net/url"
|
||||
|
||||
"github.com/tmc/langchaingo/embeddings"
|
||||
"github.com/tmc/langchaingo/llms/openai"
|
||||
"github.com/tmc/langchaingo/vectorstores/qdrant"
|
||||
)
|
||||
|
||||
func main() {
|
||||
llm, err: = openai.New()
|
||||
if err != nil {
|
||||
log.Fatal(err)
|
||||
}
|
||||
|
||||
e, err: = embeddings.NewEmbedder(llm)
|
||||
if err != nil {
|
||||
log.Fatal(err)
|
||||
}
|
||||
|
||||
url, err: = url.Parse("YOUR_QDRANT_REST_URL")
|
||||
if err != nil {
|
||||
log.Fatal(err)
|
||||
}
|
||||
|
||||
store, err: = qdrant.New(
|
||||
qdrant.WithURL(*url),
|
||||
qdrant.WithCollectionName("YOUR_COLLECTION_NAME"),
|
||||
qdrant.WithEmbedder(e),
|
||||
)
|
||||
if err != nil {
|
||||
log.Fatal(err)
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- You can find usage examples of Langchain Go [here](https://github.com/tmc/langchaingo/tree/main/examples).
|
||||
|
||||
- [Source Code](https://github.com/tmc/langchaingo/tree/main/vectorstores/qdrant)
|
||||
@@ -1,132 +0,0 @@
|
||||
---
|
||||
title: OpenAI Agents
|
||||
aliases:
|
||||
- /documentation/frameworks/swarm/
|
||||
- /articles/chatgpt-plugin/
|
||||
---
|
||||
|
||||
|
||||
# OpenAI Agents
|
||||
|
||||
[OpenAI Agents](https://github.com/openai/openai-agents-python) is a Python framework to build agentic AI apps in a lightweight, easy-to-use package with very few abstractions. It's a production-ready upgrade of the experimental framework, [Swarm](https://github.com/openai/swarm).
|
||||
|
||||
## Getting Started
|
||||
|
||||
To start using OpenAI Agents, follow these steps:
|
||||
|
||||
- Install the package
|
||||
|
||||
```bash
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
- Set up your OpenAI API key
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY="<YOUR_KEY>"
|
||||
```
|
||||
|
||||
## How It Works
|
||||
|
||||
The Agents SDK has a very small set of primitives:
|
||||
|
||||
- `Agents`, which are LLMs equipped with instructions and tools
|
||||
- `Handoffs`, which allow agents to delegate to other agents for specific tasks
|
||||
- `Guardrails`, which enable the inputs to agents to be validated
|
||||
|
||||
Used with Python, these building blocks make it easy to create real-world apps with tool-agent interactions and minimal learning curve. Plus, the SDK also includes tracing to help you debug, evaluate, and fine-tune your agent workflows.
|
||||
|
||||
## Creating Your First Agents
|
||||
|
||||
Here’s a basic example of three agents:
|
||||
|
||||
- Triage Agent: Acts as the initial point of contact. It analyzes the user's question and decides whether to route it to a specialized agent.
|
||||
- Math Tutor: A specialist agent designed to help with math-related questions.
|
||||
- History Tutor: A specialist agent focused on historical topics.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
|
||||
|
||||
math_tutor_agent = Agent(
|
||||
name="Math Tutor",
|
||||
handoff_description="Specialist agent for math questions",
|
||||
instructions="You provide help with math problems. Explain your reasoning at each step and include examples",
|
||||
)
|
||||
|
||||
history_tutor_agent = Agent(
|
||||
name="History Tutor",
|
||||
handoff_description="Specialist agent for historical questions",
|
||||
instructions="You provide assistance with historical queries. Explain important events and context clearly.",
|
||||
)
|
||||
|
||||
triage_agent = Agent(
|
||||
name="Triage Agent",
|
||||
instructions="You determine which agent to use based on the user's homework question",
|
||||
handoffs=[history_tutor_agent, math_tutor_agent],
|
||||
)
|
||||
|
||||
# Run the interaction
|
||||
result = Runner.run_sync(triage_agent, "I want some help with WW1.")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## Integrating with Qdrant
|
||||
|
||||
You can connect agents to retrieve or ingest data into a Qdrant collection. Thereby building your knowledge base. Here’s how to enable an agent to retrieve information from Qdrant.
|
||||
|
||||
Assume you have a Qdrant [collection created](https://qdrant.tech/documentation/concepts/collections/#create-a-collection) using the `"text-embedding-3-small"` model. The payload structure includes a `text` field for knowledge storage.
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
from openai import OpenAI
|
||||
from agents import Agent, function_tool
|
||||
|
||||
# Initialize clients
|
||||
openai_client = OpenAI()
|
||||
qdrant = qdrant_client.QdrantClient(host="localhost")
|
||||
|
||||
# Configuration
|
||||
EMBEDDING_MODEL = "text-embedding-3-small"
|
||||
COLLECTION_NAME = "help_center"
|
||||
LIMIT = 5
|
||||
SCORE_THRESHOLD = 0.7
|
||||
|
||||
@function_tool
|
||||
def query_qdrant(query: str) -> str:
|
||||
"""Retrieve semantically relevant content from Qdrant.
|
||||
|
||||
Args:
|
||||
query: The query to search.
|
||||
"""
|
||||
embedded_query = openai_client.embeddings.create(
|
||||
input=query,
|
||||
model=EMBEDDING_MODEL,
|
||||
).data[0].embedding
|
||||
|
||||
results = qdrant.query_points(
|
||||
collection_name=COLLECTION_NAME,
|
||||
query=embedded_query,
|
||||
limit=LIMIT,
|
||||
score_threshold=SCORE_THRESHOLD,
|
||||
).points
|
||||
|
||||
if results:
|
||||
return "\n".join([point.payload["text"] for point in results])
|
||||
else:
|
||||
return "No results found."
|
||||
|
||||
qdrant_agent = Agent(
|
||||
name="Qdrant searcher",
|
||||
handoff_description="Specialist agent for retrieving info from a Qdrant collection",
|
||||
instructions="You help find answers for user queries using Qdrant. Do not make up any info on your own.",
|
||||
tools=[query_qdrant],
|
||||
)
|
||||
```
|
||||
|
||||
Our `qdrant_agent` can now query a Qdrant collection whenever deemed necessary to answer a user query.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Agents Documentation](https://openai.github.io/openai-agents-python/)
|
||||
- [Agents Examples](https://github.com/openai/openai-agents-python/tree/main/examples)
|
||||
@@ -1,95 +0,0 @@
|
||||
---
|
||||
title: Pandas-AI
|
||||
---
|
||||
|
||||
# Pandas-AI
|
||||
|
||||
Pandas-AI is a Python library that uses a generative AI model to interpret natural language queries and translate them into Python code to interact with pandas data frames and return the final results to the user.
|
||||
|
||||
## Installation
|
||||
|
||||
```console
|
||||
pip install pandasai[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
You can begin a conversation by instantiating an `Agent` instance based on your Pandas data frame. The default Pandas-AI LLM requires an [API key](https://pandabi.ai).
|
||||
|
||||
You can find the list of all supported LLMs [here](https://docs.pandas-ai.com/en/latest/LLMs/llms/)
|
||||
|
||||
```python
|
||||
import os
|
||||
import pandas as pd
|
||||
from pandasai import Agent
|
||||
|
||||
# Sample DataFrame
|
||||
sales_by_country = pd.DataFrame(
|
||||
{
|
||||
"country": [
|
||||
"United States",
|
||||
"United Kingdom",
|
||||
"France",
|
||||
"Germany",
|
||||
"Italy",
|
||||
"Spain",
|
||||
"Canada",
|
||||
"Australia",
|
||||
"Japan",
|
||||
"China",
|
||||
],
|
||||
"sales": [5000, 3200, 2900, 4100, 2300, 2100, 2500, 2600, 4500, 7000],
|
||||
}
|
||||
)
|
||||
|
||||
os.environ["PANDASAI_API_KEY"] = "YOUR_API_KEY"
|
||||
|
||||
agent = Agent(sales_by_country)
|
||||
agent.chat("Which are the top 5 countries by sales?")
|
||||
# OUTPUT: China, United States, Japan, Germany, Australia
|
||||
```
|
||||
|
||||
## Qdrant support
|
||||
|
||||
You can train Pandas-AI to understand your data better and improve the quality of the results.
|
||||
|
||||
Qdrant can be configured as a vector store to ingest training data and retrieve semantically relevant content.
|
||||
|
||||
```python
|
||||
from pandasai.ee.vectorstores.qdrant import Qdrant
|
||||
|
||||
qdrant = Qdrant(
|
||||
collection_name="<SOME_COLLECTION>",
|
||||
embedding_model="sentence-transformers/all-MiniLM-L6-v2",
|
||||
url="http://localhost:6333",
|
||||
grpc_port=6334,
|
||||
prefer_grpc=True
|
||||
)
|
||||
|
||||
agent = Agent(df, vector_store=qdrant)
|
||||
|
||||
# Train with custom information
|
||||
agent.train(docs="The fiscal year starts in April")
|
||||
|
||||
# Train the q/a pairs of code snippets
|
||||
query = "What are the total sales for the current fiscal year?"
|
||||
response = """
|
||||
import pandas as pd
|
||||
|
||||
df = dfs[0]
|
||||
|
||||
# Calculate the total sales for the current fiscal year
|
||||
total_sales = df[df['date'] >= pd.to_datetime('today').replace(month=4, day=1)]['sales'].sum()
|
||||
result = { "type": "number", "value": total_sales }
|
||||
"""
|
||||
agent.train(queries=[query], codes=[response])
|
||||
|
||||
# # The model will use the information provided in the training to generate a response
|
||||
|
||||
```
|
||||
|
||||
## Further reading
|
||||
|
||||
- [Getting Started with Pandas-AI](https://pandasai-docs.readthedocs.io/en/latest/getting-started/)
|
||||
- [Pandas-AI Reference](https://pandasai-docs.readthedocs.io/en/latest/)
|
||||
- [Source Code](https://github.com/sinaptik-ai/pandas-ai/tree/main/extensions/ee/vectorstores/qdrant)
|
||||
@@ -1,83 +0,0 @@
|
||||
---
|
||||
title: Ragbits
|
||||
---
|
||||
|
||||
# Ragbits
|
||||
|
||||
[Ragbit](https://ragbits.deepsense.ai) is a Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. It prioritizes developer experience by providing a simple and intuitive API. It also includes a comprehensive set of tools for seamlessly building, testing, and deploying your RAG applications efficiently.
|
||||
|
||||
Qdrant is available as a vectorstore in Ragbits to ingest and search search documents from a collection.
|
||||
|
||||
## Installation
|
||||
|
||||
Install the Python package that comes bundled with the Qdrant integration.
|
||||
|
||||
```bash
|
||||
pip install ragbits
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
An example usage of Ragbits and Qdrant would look something like this:
|
||||
|
||||
The following example uses [OpenAI embeddings](https://platform.openai.com/docs/guides/embeddings) via [LiteLLM](https://www.litellm.ai).
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
from qdrant_client import AsyncQdrantClient
|
||||
|
||||
from ragbits.core.embeddings.litellm import LiteLLMEmbeddings
|
||||
from ragbits.core.vector_stores.qdrant import QdrantVectorStore
|
||||
from ragbits.document_search import DocumentSearch, SearchConfig
|
||||
from ragbits.document_search.documents.document import DocumentMeta
|
||||
|
||||
documents = [
|
||||
DocumentMeta.create_text_document_from_literal(
|
||||
"RIP boiled water. You will be mist."
|
||||
),
|
||||
DocumentMeta.create_text_document_from_literal(
|
||||
"Why programmers don't like to swim? Because they're scared of the floating points."
|
||||
),
|
||||
DocumentMeta.create_text_document_from_literal("This one is completely unrelated."),
|
||||
]
|
||||
|
||||
|
||||
async def main() -> None:
|
||||
embedder = LiteLLMEmbeddings(
|
||||
model="text-embedding-3-small",
|
||||
)
|
||||
vector_store = QdrantVectorStore(
|
||||
client=AsyncQdrantClient(url="http://localhost:6333"),
|
||||
collection_name="{collection_name}",
|
||||
)
|
||||
document_search = DocumentSearch(
|
||||
embedder=embedder,
|
||||
vector_store=vector_store,
|
||||
)
|
||||
|
||||
await document_search.ingest(documents)
|
||||
|
||||
all_documents = await vector_store.list()
|
||||
print([doc.metadata["content"] for doc in all_documents])
|
||||
|
||||
query = "I write computer software. Tell me something."
|
||||
vector_store_kwargs = {
|
||||
"k": 1,
|
||||
"max_distance": None,
|
||||
}
|
||||
results = await document_search.search(
|
||||
query,
|
||||
config=SearchConfig(vector_store_kwargs=vector_store_kwargs),
|
||||
)
|
||||
|
||||
print(f"Documents similar to: {query}")
|
||||
print([element.get_key() for element in results])
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- Ragbits [Documentation](http://ragbits.deepsense.ai)
|
||||
- [Source Code](https://github.com/deepsense-ai/ragbits)
|
||||
@@ -1,94 +0,0 @@
|
||||
---
|
||||
title: Solon
|
||||
---
|
||||
|
||||
# Solon
|
||||
|
||||
[Solon](https://solon.noear.org) is a lightweight, high-performance Java enterprise framework designed for efficient, eco-friendly development. It enhances concurrency, reduces memory usage, speeds up startup, minimizes packaging size, and supports Java 8 to Java 23, offering a flexible alternative to Spring.
|
||||
|
||||
Qdrant is available as a component in Solon-AI for efficient vector indexing and retrievals.
|
||||
|
||||
## Installation
|
||||
|
||||
```xml
|
||||
<dependency>
|
||||
<groupId>org.noear</groupId>
|
||||
<artifactId>solon-ai-repo-qdrant</artifactId>
|
||||
</dependency>
|
||||
```
|
||||
|
||||
This is the main extension plugin for **solon-ai**, which provides the `QdrantRepository` knowledge base.
|
||||
|
||||
## Configuration
|
||||
|
||||
When using `QdrantRepository`, an embedding model needs to be configured.
|
||||
|
||||
```yaml
|
||||
solon.ai.embed:
|
||||
bgem3:
|
||||
apiUrl: "http://127.0.0.1:11434/api/embed"
|
||||
provider: "ollama"
|
||||
model: "bge-m3:latest"
|
||||
|
||||
solon.ai.repo:
|
||||
qdrant:
|
||||
host: "localhost"
|
||||
port: 6334
|
||||
useSsl: false
|
||||
```
|
||||
|
||||
You can now instantiate the embedding model and Qdrant.
|
||||
|
||||
```java
|
||||
@Configuration
|
||||
public class DemoConfig {
|
||||
|
||||
// Create the embedding model
|
||||
@Bean
|
||||
public EmbeddingModel embeddingModel(@Inject("${solon.ai.embed.bgem3}") EmbeddingConfig config) {
|
||||
return EmbeddingModel.of(config).build();
|
||||
}
|
||||
|
||||
// Configure the QdrantClient using QdrantGrpcClient
|
||||
@BindProps(prefix = "solon.ai.repo.qdrant")
|
||||
@Bean
|
||||
public QdrantClient qdrantClient(@Value("${solon.ai.repo.qdrant.host}") String host,
|
||||
@Value("${solon.ai.repo.qdrant.port}") int port,
|
||||
@Value("${solon.ai.repo.qdrant.useSsl}") boolean useSsl) {
|
||||
return new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder(host, port, useSsl).build()
|
||||
);
|
||||
}
|
||||
|
||||
// Initialize the Qdrant knowledge base
|
||||
@Bean
|
||||
public QdrantRepository repository(EmbeddingModel embeddingModel, QdrantClient client) {
|
||||
return new QdrantRepository(embeddingModel, client);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```java
|
||||
@Component
|
||||
public class DemoService {
|
||||
@Inject
|
||||
private QdrantRepository repository;
|
||||
|
||||
// Add documents to the repository
|
||||
public void addDocument(List<Document> docs) {
|
||||
repository.insert(docs);
|
||||
}
|
||||
|
||||
// Search for documents based on a query
|
||||
public List<Document> findDocument(String query) {
|
||||
return repository.search(query);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- Solon [Documentation](https://solon.noear.org).
|
||||
- Solon [Source](https://github.com/opensolon/solon)
|
||||
@@ -1,87 +0,0 @@
|
||||
---
|
||||
title: Superduper
|
||||
---
|
||||
|
||||
# Superduper
|
||||
|
||||
[Superduper](https://superduper.io/) is a framework for building flexible, compositional AI applications which may be applied directly to databases using a declarative programming model. These applications declare and maintain a desired state of the database, and use the database directly to store outputs of AI components, meta-data about the components and data pertaining to the state of the system.
|
||||
|
||||
Qdrant is available as a vector search provider in Superduper.
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
pip install superduper-framework
|
||||
```
|
||||
|
||||
## Setup
|
||||
|
||||
- To use Qdrant for vector search layer, create a `settings.yaml` file with the following config.
|
||||
|
||||
> settings.yaml
|
||||
```yaml
|
||||
cluster:
|
||||
vector_search:
|
||||
type: qdrant
|
||||
|
||||
vector_search_kwargs:
|
||||
url: "http://localhost:6333"
|
||||
api_key: "<YOUR_API_KEY>
|
||||
# Supports all parameters of qdrant_client.QdrantClient
|
||||
```
|
||||
|
||||
- Set the `SUPERDUPER_CONFIG` env value to the path of the config file.
|
||||
|
||||
```bash
|
||||
export SUPERDUPER_CONFIG=path/to/settings.yaml
|
||||
```
|
||||
|
||||
That's all. You can now use Superduper backed by Qdrant.
|
||||
|
||||
## Example
|
||||
|
||||
Here's an example to run vector search using the configured Qdrant index.
|
||||
|
||||
```python
|
||||
import json
|
||||
import requests
|
||||
from superduper import superduper, Document
|
||||
from superduper.ext.sentence_transformers import SentenceTransformer
|
||||
|
||||
r = requests.get('https://superduperdb-public-demo.s3.amazonaws.com/text.json')
|
||||
|
||||
with open('text.json', 'wb') as f:
|
||||
f.write(r.content)
|
||||
|
||||
with open('text.json', 'r') as f:
|
||||
data = json.load(f)
|
||||
|
||||
db = superduper('mongomock://test')
|
||||
|
||||
_ = db['documents'].insert_many([Document({'txt': txt}) for txt in data]).execute()
|
||||
|
||||
model = SentenceTransformer(
|
||||
identifier="test",
|
||||
predict_kwargs={"show_progress_bar": True},
|
||||
model="all-MiniLM-L6-v2",
|
||||
device="cpu",
|
||||
postprocess=lambda x: x.tolist(),
|
||||
)
|
||||
|
||||
vector_index = model.to_vector_index(select=db['documents'].find(), key='txt')
|
||||
|
||||
db.apply(vector_index)
|
||||
|
||||
query = db['documents'].like({'txt': 'Tell me about vector-search'}, vector_index=vector_index.identifier, n=3).find()
|
||||
cursor = query.execute()
|
||||
|
||||
for r in cursor:
|
||||
print('=' * 100)
|
||||
print(r.unpack()['txt'])
|
||||
print('=' * 100)
|
||||
```
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- Superduper [Intro](https://docs.superduper.io/docs/intro)
|
||||
- Superduper [Source](https://github.com/superduper-io/superduper)
|
||||
|
Before Width: | Height: | Size: 22 KiB |
@@ -9,17 +9,13 @@ partition: build
|
||||
| Platform | Description |
|
||||
| ------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
|
||||
| [Apify](/documentation/platforms/apify/) | Platform to build web scrapers and automate web browser tasks. |
|
||||
| [Bubble](/documentation/platforms/bubble/) | Development platform for application development with a no-code interface |
|
||||
| [BuildShip](/documentation/platforms/buildship/) | Low-code visual builder to create APIs, scheduled jobs, and backend workflows. |
|
||||
| [DocsGPT](/documentation/platforms/docsgpt/) | Tool for ingesting documentation sources and enabling conversations and queries. |
|
||||
| [Keboola](/documentation/platforms/keboola/) | Data operations platform that unifies data sources, transformations, and ML deployments. |
|
||||
| [Kotaemon](/documentation/platforms/kotaemon/) | Open-source & customizable RAG UI for chatting with your documents. |
|
||||
| [Make](/documentation/platforms/make/) | Cloud platform to build low-code workflows by integrating various software applications. |
|
||||
| [Mulesoft Anypoint](/documentation/platforms/mulesoft/) | Integration platform to connect applications, data, and devices across environments. |
|
||||
| [N8N](/documentation/platforms/n8n/) | Platform for node-based, low-code workflow automation. |
|
||||
| [Pipedream](/documentation/platforms/pipedream/) | Platform for connecting apps and developing event-driven automation. |
|
||||
| [Portable.io](/documentation/platforms/portable/) | Cloud platform for developing and deploying ELT transformations. |
|
||||
| [PrivateGPT](/documentation/platforms/privategpt/) | Tool to ask questions about your documents using local LLMs emphasising privacy. |
|
||||
| [Rivet](/documentation/platforms/rivet/) | A visual programming environment for building AI agents with LLMs. |
|
||||
| [ToolJet](/documentation/platforms/tooljet/) | A low-code platform for business apps that connect to DBs, cloud storages and more. |
|
||||
| [Vectorize](/documentation/platforms/vectorize/) | Platform to automate data extraction, RAG evaluation, deploy RAG pipelines. |
|
||||
|
||||
@@ -1,36 +0,0 @@
|
||||
---
|
||||
title: Bubble
|
||||
aliases: [ ../frameworks/bubble/ ]
|
||||
---
|
||||
|
||||
# Bubble
|
||||
|
||||
[Bubble](https://bubble.io/) is a software development platform that enables anyone to build and launch fully functional web applications without writing code.
|
||||
|
||||
You can use the [Qdrant Bubble plugin](https://bubble.io/plugin/qdrant-1716804374179x344999530386685950) to interface with Qdrant in your workflows.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
1. A Qdrant instance to connect to. You can get a free cloud instance at [cloud.qdrant.io](https://cloud.qdrant.io/).
|
||||
2. An account at [Bubble.io](https://bubble.io/) and an app set up.
|
||||
|
||||
## Setting up the plugin
|
||||
|
||||
Navigate to your app's workflows. Select `"Install more plugins actions"`.
|
||||
|
||||

|
||||
|
||||
You can now search for the Qdrant plugin and install it. Ensure all the categories are selected to perform a full search.
|
||||
|
||||

|
||||
|
||||
The Qdrant plugin can now be found in the installed plugins section of your workflow. Enter the API key of your Qdrant instance for authentication.
|
||||
|
||||

|
||||
|
||||
The plugin provides actions for upserting, searching, updating and deleting points from your Qdrant collection with dynamic and static values from your Bubble workflow.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Bubble Academy](https://bubble.io/academy).
|
||||
- [Bubble Manual](https://manual.bubble.io/)
|
||||
@@ -1,28 +0,0 @@
|
||||
---
|
||||
title: DocsGPT
|
||||
aliases: [ ../frameworks/docsgpt/ ]
|
||||
---
|
||||
|
||||
# DocsGPT
|
||||
|
||||
[DocsGPT](https://docsgpt.arc53.com/) is an open-source documentation assistant that enables you to build conversational user experiences on top of your data.
|
||||
|
||||
Qdrant is supported as a vectorstore in DocsGPT to ingest and semantically retrieve documents.
|
||||
|
||||
## Configuration
|
||||
|
||||
Learn how to setup DocsGPT in their [Quickstart guide](https://docs.docsgpt.cloud/quickstart).
|
||||
|
||||
You can configure DocsGPT with environment variables in a `.env` file.
|
||||
|
||||
To configure DocsGPT to use Qdrant as the vector store, set `VECTOR_STORE` to `"qdrant"`.
|
||||
|
||||
```bash
|
||||
echo "VECTOR_STORE=qdrant" >> .env
|
||||
```
|
||||
|
||||
DocsGPT includes a list of the Qdrant configuration options that you can set as environment variables [here](https://github.com/arc53/DocsGPT/blob/00dfb07b15602319bddb95089e3dab05fac56240/application/core/settings.py#L46-L59).
|
||||
|
||||
## Further reading
|
||||
|
||||
- [DocsGPT Reference](https://github.com/arc53/DocsGPT)
|
||||
@@ -1,34 +0,0 @@
|
||||
---
|
||||
title: Portable.io
|
||||
aliases: [ ../frameworks/portable/ ]
|
||||
---
|
||||
|
||||
# Portable
|
||||
|
||||
[Portable](https://portable.io/) is an ELT platform that builds connectors on-demand for data teams. It enables connecting applications to your data warehouse with no code.
|
||||
|
||||
You can avail the [Qdrant connector](https://portable.io/connectors/qdrant) to build data pipelines from your collections.
|
||||
|
||||

|
||||
|
||||
## Prerequisites
|
||||
|
||||
1. A Qdrant instance to connect to. You can get a free cloud instance at [cloud.qdrant.io](https://cloud.qdrant.io/).
|
||||
2. A [Portable account](https://app.portable.io/).
|
||||
|
||||
## Setting up the connector
|
||||
|
||||
Navigate to the Portable dashboard. Search for `"Qdrant"` in the sources section.
|
||||
|
||||

|
||||
|
||||
Configure the connector with your Qdrant instance credentials.
|
||||
|
||||

|
||||
|
||||
You can now build your flows using data from Qdrant by selecting a [destination](https://app.portable.io/destinations) and scheduling it.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Portable API Reference](https://developer.portable.io/api-reference/introduction).
|
||||
- [Portable Academy](https://portable.io/learn)
|
||||
@@ -1,34 +0,0 @@
|
||||
---
|
||||
title: Ironclad Rivet
|
||||
aliases: [ ../frameworks/rivet/ ]
|
||||
---
|
||||
|
||||
# Ironclad Rivet
|
||||
|
||||
[Rivet](https://rivet.ironcladapp.com/) is an Integrated Development Environment (IDE) and library designed for creating AI agents using a visual, graph-based interface.
|
||||
|
||||
Qdrant is available as a [plugin](https://github.com/qdrant/rivet-plugin-qdrant) for building vector-search powered workflows in Rivet.
|
||||
|
||||
## Installation
|
||||
|
||||
- Open the plugins overlay at the top of the screen.
|
||||
- Search for the official Qdrant plugin.
|
||||
- Click the "Add" button to install it in your current project.
|
||||
|
||||

|
||||
|
||||
## Setting up the connection
|
||||
|
||||
You can configure your Qdrant instance credentials in the Rivet settings after installing the plugin.
|
||||
|
||||

|
||||
|
||||
Once you've configured your credentials, you can right-click on your workspace to add nodes from the plugin and get building!
|
||||
|
||||

|
||||
|
||||
## Further Reading
|
||||
|
||||
- Rivet [Tutorial](https://rivet.ironcladapp.com/docs/tutorial).
|
||||
- Rivet [Documentation](https://rivet.ironcladapp.com/docs).
|
||||
- Plugin [Source Code](https://github.com/qdrant/rivet-plugin-qdrant)
|
||||
@@ -13,7 +13,7 @@ title: ToolJet
|
||||
|
||||
## Setting Up
|
||||
|
||||
- Search for the Qdrant plugin in the Tooljet [plugins marketplace](https://docs.tooljet.com/docs/marketplace/plugins/marketplace-plugin-qdrant).
|
||||
- Search for the Qdrant plugin in the Tooljet [plugins marketplace](https://docs.tooljet.ai/docs/marketplace/plugins/marketplace-plugin-qdrant/).
|
||||
|
||||
- Set up the connection to Qdrant using your instance credentials.
|
||||
|
||||
@@ -48,7 +48,5 @@ You can interface with the Qdrant instance using the following Tooljet operation
|
||||
## Further Reading
|
||||
|
||||
- [ToolJet Documentation](https://docs.tooljet.com/docs/).
|
||||
<!---
|
||||
- [ToolJet Qdrant Plugin](https://docs.tooljet.com/docs/marketplace/plugins/marketplace-plugin-qdrant/).
|
||||
-->
|
||||
- [ToolJet Qdrant Plugin](https://docs.tooljet.ai/docs/marketplace/plugins/marketplace-plugin-qdrant/).
|
||||
|
||||
|
||||
|
Before Width: | Height: | Size: 80 KiB |
|
Before Width: | Height: | Size: 151 KiB |
|
Before Width: | Height: | Size: 153 KiB |
|
Before Width: | Height: | Size: 324 KiB |
|
Before Width: | Height: | Size: 374 KiB |
|
Before Width: | Height: | Size: 371 KiB |
|
Before Width: | Height: | Size: 109 KiB |
|
Before Width: | Height: | Size: 66 KiB |
|
Before Width: | Height: | Size: 68 KiB |
|
Before Width: | Height: | Size: 326 KiB |
|
Before Width: | Height: | Size: 166 KiB |
|
Before Width: | Height: | Size: 302 KiB |
|
Before Width: | Height: | Size: 122 KiB |
|
Before Width: | Height: | Size: 105 KiB |
|
Before Width: | Height: | Size: 109 KiB |
|
Before Width: | Height: | Size: 136 KiB |