docs: Clean up integration docs (#1840)

* docs: Clean up integration docs

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: More cleanup

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: Deleted image artifacts

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: Updated redirects

Signed-off-by: Anush008 <anushshetty90@gmail.com>

---------

Signed-off-by: Anush008 <anushshetty90@gmail.com>
This commit is contained in:
Anush
2025-08-08 22:24:51 +05:30
committed by GitHub
parent 03455c7e46
commit 9968ce4842
43 changed files with 125 additions and 1209 deletions
@@ -11,14 +11,11 @@ aliases: ["/documentation/frameworks/memgpt/"]
| ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| [AutoGen](/documentation/frameworks/autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
| [Camel](/documentation/frameworks/camel/) | Framework to build and use LLM-based agents for real-world task solving |
| [Canopy](/documentation/frameworks/canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
| [Cheshire Cat](/documentation/frameworks/cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
| [CrewAI](/documentation/frameworks/crewai/) | CrewAI is a framework to build automated workflows using multiple AI agents that perform complex tasks. |
| [Dagster](/documentation/frameworks/dagster/) | Python framework for data orchestration with integrated lineage, observability. |
| [DeepEval](/documentation/frameworks/deepeval/) | Python framework for testing large language model systems. |
| [DocArray](/documentation/frameworks/docarray/) | Python library for managing data in multi-modal AI applications. |
| [DSPy](/documentation/frameworks/dspy/) | Framework for algorithmically optimizing LM prompts and weights. |
| [dsRAG](/documentation/frameworks/dsrag/) | High-performance Python retrieval engine for unstructured data. |
| [Dynamiq](/documentation/frameworks/dynamiq/) | Dynamiq is all-in-one Gen AI framework, designed to streamline the development of AI-powered applications. |
| [Feast](/documentation/frameworks/feast/) | Open-source feature store to operate production ML systems at scale as a set of features. |
| [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. |
@@ -27,7 +24,6 @@ aliases: ["/documentation/frameworks/memgpt/"]
| [HoneyHive](/documentation/frameworks/honeyhive/) | AI observability and evaluation platform that provides tracing and monitoring tools for GenAI pipelines. |
| [Lakechain](/documentation/frameworks/lakechain/) | Python framework for deploying document processing pipelines on AWS using infrastructure-as-code. |
| [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. |
| [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. |
| [Langchain4j](/documentation/frameworks/langchain4j/) | Java framework for building context-aware, reasoning applications using LLMs. |
| [LangGraph](/documentation/frameworks/langgraph/) | Python, Javascript libraries for building stateful, multi-actor applications. |
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
@@ -36,15 +32,10 @@ aliases: ["/documentation/frameworks/memgpt/"]
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
| [Neo4j GraphRAG](/documentation/frameworks/neo4j-graphrag/) | Package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. |
| [NLWeb](/documentation/frameworks/nlweb/) | A framework to turn websites into chat-ready data using schema.org and associated data formats. |
| [OpenAI Agents](/documentation/frameworks/openai-agents/) | Python framework for managing multiple AI agents that can work together. |
| [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language |
| [Ragbits](/documentation/frameworks/ragbits/) | Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. |
| [Rig-rs](/documentation/frameworks/rig-rs/) | Rust library for building scalable, modular, and ergonomic LLM-powered applications. |
| [Semantic Router](/documentation/frameworks/semantic-router/) | Python library to build a decision-making layer for AI applications using vector search. |
| [SmolAgents](/documentation/frameworks/smolagents/) | Barebones library for agents. Agents write python code to call tools and orchestrate other agent. |
| [Solon](/documentation/frameworks/solon/) | A lightweight, high-performance Java enterprise framework |
| [Spring AI](/documentation/frameworks/spring-ai/) | Java AI framework for building with Spring design principles such as portability and modular design. |
| [Superduper](/documentation/frameworks/superduper/) | Framework for building flexible, compositional AI apps which may be applied directly to databases. |
| [Sycamore](/documentation/frameworks/sycamore/) | Document processing engine for ETL, RAG, LLM-based applications, and analytics on unstructured data. |
| [Testcontainers](/documentation/frameworks/testcontainers/) | Framework for providing throwaway, lightweight instances of systems for testing |
| [txtai](/documentation/frameworks/txtai/) | Python library for semantic search, LLM orchestration and language model workflows. |
@@ -1,90 +0,0 @@
---
title: Pinecone Canopy
---
# Pinecone Canopy
[Canopy](https://github.com/pinecone-io/canopy) is an open-source framework and context engine to build chat assistants at scale.
Qdrant is supported as a knowledge base within Canopy for context retrieval and augmented generation.
## Usage
Install the SDK with the Qdrant extra as described in the [Canopy README](https://github.com/pinecone-io/canopy?tab=readme-ov-file#extras).
```bash
pip install canopy-sdk[qdrant]
```
### Creating a knowledge base
```python
from canopy.knowledge_base import QdrantKnowledgeBase
kb = QdrantKnowledgeBase(collection_name="<YOUR_COLLECTION_NAME>")
```
<aside role="status">The constructor accepts additional <a href="https://github.com/qdrant/qdrant-client/blob/eda201a1dbf1bbc67415f8437a5619f6f83e8ac6/qdrant_client/qdrant_client.py#L36-L61">options</a> to customize your connection to Qdrant.</aside>
To create a new Qdrant collection and connect it to the knowledge base, use the `create_canopy_collection` method:
```python
kb.create_canopy_collection()
```
You can always verify the connection to the collection with the `verify_index_connection` method:
```python
kb.verify_index_connection()
```
Learn more about customizing the knowledge base and its inner components [in the Canopy library](https://github.com/pinecone-io/canopy/blob/main/docs/library.md#understanding-knowledgebase-workings).
### Adding data to the knowledge base
To insert data into the knowledge base, you can create a list of documents and use the `upsert` method:
```python
from canopy.models.data_models import Document
documents = [
Document(
id="1",
text="U2 are an Irish rock band from Dublin, formed in 1976.",
source="https://en.wikipedia.org/wiki/U2",
),
Document(
id="2",
text="Arctic Monkeys are an English rock band formed in Sheffield in 2002.",
source="https://en.wikipedia.org/wiki/Arctic_Monkeys",
metadata={"my-key": "my-value"},
),
]
kb.upsert(documents)
```
### Querying the knowledge base
You can query the knowledge base with the `query` method to find the most similar documents to a given text:
```python
from canopy.models.data_models import Query
kb.query(
[
Query(text="Arctic Monkeys music genre"),
Query(
text="U2 music genre",
top_k=10,
metadata_filter={"key": "my-key", "match": {"value": "my-value"}},
),
]
)
```
## Further Reading
- [Introduction to Canopy](https://www.pinecone.io/blog/canopy-rag-framework/)
- [Canopy library reference](https://github.com/pinecone-io/canopy/blob/main/docs/library.md)
- [Source Code](https://github.com/pinecone-io/canopy/tree/main/src/canopy/knowledge_base/qdrant)
@@ -1,22 +0,0 @@
---
title: DocArray
aliases: [ ../integrations/docarray/ ]
---
# DocArray
You can use Qdrant natively in DocArray, where Qdrant serves as a high-performance document store to enable scalable vector search.
DocArray is a library from Jina AI for nested, unstructured data in transit, including text, image, audio, video, 3D mesh, etc.
It allows deep-learning engineers to efficiently process, embed, search, recommend, store, and transfer the data with a Pythonic API.
To install DocArray with Qdrant support, please do
```bash
pip install "docarray[qdrant]"
```
## Further Reading
- [DocArray documentations](https://docarray.jina.ai/advanced/document-store/qdrant/).
- [Source Code](https://github.com/docarray/docarray/blob/main/docarray/index/backends/qdrant.py)
@@ -1,50 +0,0 @@
---
title: dsRAG
---
# dsRAG
[dsRAG](https://github.com/D-Star-AI/dsRAG) is a retrieval engine for unstructured data. It is especially good at handling challenging queries over dense text, like financial reports, legal documents, and academic papers. dsRAG achieves substantially higher accuracy than vanilla RAG baselines on complex open-book question answering tasks
You can use the Qdrant connector in dsRAG to add and semantically retrieve documents from your collections.
## Usage Example
```python
from dsrag.database.vector import QdrantVectorDB
import numpy as np
from qdrant_clien import models
db = QdrantVectorDB(kb_id=self.kb_id, url="http://localhost:6334", prefer_grpc=True)
vectors = [np.array([1, 0]), np.array([0, 1])]
# You can use any document loaders available with dsRAG
# We'll use literals for demonstration
documents = [
{
"doc_id": "1",
"chunk_index": 0,
"chunk_header": "Header1",
"chunk_text": "Text1",
},
{
"doc_id": "2",
"chunk_index": 1,
"chunk_header": "Header2",
"chunk_text": "Text2",
},
]
db.add_vectors(vectors, documents)
metadata_filter = models.Filter(
must=[models.FieldCondition(key="doc_id", match=models.MatchValue(value="1"))]
)
db.search(query_vector, top_k=4, metadata_filter=metadata_filter)
```
## Further Reading
- [dsRAG Source](https://github.com/D-Star-AI/dsRAG).
- [dsRAG Examples](https://github.com/D-Star-AI/dsRAG/tree/main/examples)
@@ -59,4 +59,4 @@ feature_values = feature_store.retrieve_online_documents(
## 📚 Further Reading
- [Feast Documentation](http://docs.feast.dev/)
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/)
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/qdrant_online_store)
@@ -24,26 +24,23 @@ To use this plugin, specify it when you call `configureGenkit()`:
```js
import { qdrant } from 'genkitx-qdrant';
import { textEmbeddingGecko } from '@genkit-ai/vertexai';
export default configureGenkit({
plugins: [
qdrant([
{
clientParams: {
host: 'localhost',
port: 6333,
},
collectionName: 'some-collection',
embedder: textEmbeddingGecko,
},
]),
],
// ...
const ai = genkit({
plugins: [
qdrant([
{
embedder: googleAI.embedder('text-embedding-004'),
collectionName: 'collectionName',
clientParams: {
url: 'http://localhost:6333',
}
}
]),
],
});
```
You'll need to specify a collection name, the embedding model you want to use and the Qdrant client parameters. In
You'll need to specify a collection name, the embedding model you want to use and the Qdrant client parameters. In
addition, there are a few optional parameters:
- `embedderOptions`: Additional options to pass options to the embedder:
@@ -64,7 +61,13 @@ addition, there are a few optional parameters:
metadataPayloadKey: 'metadata';
```
- `collectionCreateOptions`: [Additional options](/documentation/concepts/collections/#create-a-collection/) when creating the Qdrant collection.
- `dataTypePayloadKey`: Name of the payload filed with the document datatype. Defaults to "_content_type".
```js
dataTypePayloadKey: '_datatype';
```
- `collectionCreateOptions`: [Additional options](<(https://qdrant.tech/documentation/concepts/collections/#create-a-collection)>) when creating the Qdrant collection.
## Usage
@@ -72,36 +75,25 @@ Import retriever and indexer references like so:
```js
import { qdrantIndexerRef, qdrantRetrieverRef } from 'genkitx-qdrant';
import { Document, index, retrieve } from '@genkit-ai/ai/retriever';
```
Then, pass the references to `retrieve()` and `index()`:
Then, pass their references to `retrieve()` and `index()`:
```js
// To specify an indexer:
export const qdrantIndexer = qdrantIndexerRef({
collectionName: 'some-collection',
displayName: 'Some Collection indexer',
});
await index({ indexer: qdrantIndexer, documents });
// To export an indexer reference:
export const qdrantIndexer = qdrantIndexerRef('collectionName', 'displayName');
```
```js
// To specify a retriever:
export const qdrantRetriever = qdrantRetrieverRef({
collectionName: 'some-collection',
displayName: 'Some Collection Retriever',
});
let docs = await retrieve({ retriever: qdrantRetriever, query });
// To export a retriever reference:
export const qdrantRetriever = qdrantRetrieverRef('collectionName', 'displayName');
```
You can refer to [Retrieval-augmented generation](https://firebase.google.com/docs/genkit/rag) for a general
You can refer to [Retrieval-augmented generation](https://genkit.dev/docs/rag/) for a general
discussion on indexers and retrievers.
## Further Reading
- [Introduction to Genkit](https://firebase.google.com/docs/genkit)
- [Genkit Documentation](https://firebase.google.com/docs/genkit/get-started)
- [Introduction to Genkit](https://genkit.dev/)
- [Genkit Documentation](https://genkit.dev/docs/get-started/)
- [Source Code](https://github.com/qdrant/qdrant-genkit)
@@ -1,71 +0,0 @@
---
title: Langchain Go
---
# Langchain Go
[Langchain Go](https://tmc.github.io/langchaingo/docs/) is a framework for developing data-aware applications powered by language models in Go.
You can use Qdrant as a vector store in Langchain Go.
## Setup
Install the `langchain-go` project dependency
```bash
go get -u github.com/tmc/langchaingo
```
## Usage
Before you use the following code sample, customize the following values for your configuration:
- `YOUR_QDRANT_REST_URL`: If you've set up Qdrant using the [Quick Start](/documentation/quick-start/) guide,
set this value to `http://localhost:6333`.
- `YOUR_COLLECTION_NAME`: Use our [Collections](/documentation/concepts/collections/) guide to create or
list collections.
```go
package main
import (
"log"
"net/url"
"github.com/tmc/langchaingo/embeddings"
"github.com/tmc/langchaingo/llms/openai"
"github.com/tmc/langchaingo/vectorstores/qdrant"
)
func main() {
llm, err: = openai.New()
if err != nil {
log.Fatal(err)
}
e, err: = embeddings.NewEmbedder(llm)
if err != nil {
log.Fatal(err)
}
url, err: = url.Parse("YOUR_QDRANT_REST_URL")
if err != nil {
log.Fatal(err)
}
store, err: = qdrant.New(
qdrant.WithURL(*url),
qdrant.WithCollectionName("YOUR_COLLECTION_NAME"),
qdrant.WithEmbedder(e),
)
if err != nil {
log.Fatal(err)
}
}
```
## Further Reading
- You can find usage examples of Langchain Go [here](https://github.com/tmc/langchaingo/tree/main/examples).
- [Source Code](https://github.com/tmc/langchaingo/tree/main/vectorstores/qdrant)
@@ -1,132 +0,0 @@
---
title: OpenAI Agents
aliases:
- /documentation/frameworks/swarm/
- /articles/chatgpt-plugin/
---
# OpenAI Agents
[OpenAI Agents](https://github.com/openai/openai-agents-python) is a Python framework to build agentic AI apps in a lightweight, easy-to-use package with very few abstractions. It's a production-ready upgrade of the experimental framework, [Swarm](https://github.com/openai/swarm).
## Getting Started
To start using OpenAI Agents, follow these steps:
- Install the package
```bash
pip install openai-agents
```
- Set up your OpenAI API key
```bash
export OPENAI_API_KEY="<YOUR_KEY>"
```
## How It Works
The Agents SDK has a very small set of primitives:
- `Agents`, which are LLMs equipped with instructions and tools
- `Handoffs`, which allow agents to delegate to other agents for specific tasks
- `Guardrails`, which enable the inputs to agents to be validated
Used with Python, these building blocks make it easy to create real-world apps with tool-agent interactions and minimal learning curve. Plus, the SDK also includes tracing to help you debug, evaluate, and fine-tune your agent workflows.
## Creating Your First Agents
Here’s a basic example of three agents:
- Triage Agent: Acts as the initial point of contact. It analyzes the user's question and decides whether to route it to a specialized agent.
- Math Tutor: A specialist agent designed to help with math-related questions.
- History Tutor: A specialist agent focused on historical topics.
```python
from agents import Agent, Runner
math_tutor_agent = Agent(
name="Math Tutor",
handoff_description="Specialist agent for math questions",
instructions="You provide help with math problems. Explain your reasoning at each step and include examples",
)
history_tutor_agent = Agent(
name="History Tutor",
handoff_description="Specialist agent for historical questions",
instructions="You provide assistance with historical queries. Explain important events and context clearly.",
)
triage_agent = Agent(
name="Triage Agent",
instructions="You determine which agent to use based on the user's homework question",
handoffs=[history_tutor_agent, math_tutor_agent],
)
# Run the interaction
result = Runner.run_sync(triage_agent, "I want some help with WW1.")
print(result.final_output)
```
## Integrating with Qdrant
You can connect agents to retrieve or ingest data into a Qdrant collection. Thereby building your knowledge base. Here’s how to enable an agent to retrieve information from Qdrant.
Assume you have a Qdrant [collection created](https://qdrant.tech/documentation/concepts/collections/#create-a-collection) using the `"text-embedding-3-small"` model. The payload structure includes a `text` field for knowledge storage.
```python
import qdrant_client
from openai import OpenAI
from agents import Agent, function_tool
# Initialize clients
openai_client = OpenAI()
qdrant = qdrant_client.QdrantClient(host="localhost")
# Configuration
EMBEDDING_MODEL = "text-embedding-3-small"
COLLECTION_NAME = "help_center"
LIMIT = 5
SCORE_THRESHOLD = 0.7
@function_tool
def query_qdrant(query: str) -> str:
"""Retrieve semantically relevant content from Qdrant.
Args:
query: The query to search.
"""
embedded_query = openai_client.embeddings.create(
input=query,
model=EMBEDDING_MODEL,
).data[0].embedding
results = qdrant.query_points(
collection_name=COLLECTION_NAME,
query=embedded_query,
limit=LIMIT,
score_threshold=SCORE_THRESHOLD,
).points
if results:
return "\n".join([point.payload["text"] for point in results])
else:
return "No results found."
qdrant_agent = Agent(
name="Qdrant searcher",
handoff_description="Specialist agent for retrieving info from a Qdrant collection",
instructions="You help find answers for user queries using Qdrant. Do not make up any info on your own.",
tools=[query_qdrant],
)
```
Our `qdrant_agent` can now query a Qdrant collection whenever deemed necessary to answer a user query.
## Further Reading
- [Agents Documentation](https://openai.github.io/openai-agents-python/)
- [Agents Examples](https://github.com/openai/openai-agents-python/tree/main/examples)
@@ -1,95 +0,0 @@
---
title: Pandas-AI
---
# Pandas-AI
Pandas-AI is a Python library that uses a generative AI model to interpret natural language queries and translate them into Python code to interact with pandas data frames and return the final results to the user.
## Installation
```console
pip install pandasai[qdrant]
```
## Usage
You can begin a conversation by instantiating an `Agent` instance based on your Pandas data frame. The default Pandas-AI LLM requires an [API key](https://pandabi.ai).
You can find the list of all supported LLMs [here](https://docs.pandas-ai.com/en/latest/LLMs/llms/)
```python
import os
import pandas as pd
from pandasai import Agent
# Sample DataFrame
sales_by_country = pd.DataFrame(
{
"country": [
"United States",
"United Kingdom",
"France",
"Germany",
"Italy",
"Spain",
"Canada",
"Australia",
"Japan",
"China",
],
"sales": [5000, 3200, 2900, 4100, 2300, 2100, 2500, 2600, 4500, 7000],
}
)
os.environ["PANDASAI_API_KEY"] = "YOUR_API_KEY"
agent = Agent(sales_by_country)
agent.chat("Which are the top 5 countries by sales?")
# OUTPUT: China, United States, Japan, Germany, Australia
```
## Qdrant support
You can train Pandas-AI to understand your data better and improve the quality of the results.
Qdrant can be configured as a vector store to ingest training data and retrieve semantically relevant content.
```python
from pandasai.ee.vectorstores.qdrant import Qdrant
qdrant = Qdrant(
collection_name="<SOME_COLLECTION>",
embedding_model="sentence-transformers/all-MiniLM-L6-v2",
url="http://localhost:6333",
grpc_port=6334,
prefer_grpc=True
)
agent = Agent(df, vector_store=qdrant)
# Train with custom information
agent.train(docs="The fiscal year starts in April")
# Train the q/a pairs of code snippets
query = "What are the total sales for the current fiscal year?"
response = """
import pandas as pd
df = dfs[0]
# Calculate the total sales for the current fiscal year
total_sales = df[df['date'] >= pd.to_datetime('today').replace(month=4, day=1)]['sales'].sum()
result = { "type": "number", "value": total_sales }
"""
agent.train(queries=[query], codes=[response])
# # The model will use the information provided in the training to generate a response
```
## Further reading
- [Getting Started with Pandas-AI](https://pandasai-docs.readthedocs.io/en/latest/getting-started/)
- [Pandas-AI Reference](https://pandasai-docs.readthedocs.io/en/latest/)
- [Source Code](https://github.com/sinaptik-ai/pandas-ai/tree/main/extensions/ee/vectorstores/qdrant)
@@ -1,83 +0,0 @@
---
title: Ragbits
---
# Ragbits
[Ragbit](https://ragbits.deepsense.ai) is a Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. It prioritizes developer experience by providing a simple and intuitive API. It also includes a comprehensive set of tools for seamlessly building, testing, and deploying your RAG applications efficiently.
Qdrant is available as a vectorstore in Ragbits to ingest and search search documents from a collection.
## Installation
Install the Python package that comes bundled with the Qdrant integration.
```bash
pip install ragbits
```
## Usage
An example usage of Ragbits and Qdrant would look something like this:
The following example uses [OpenAI embeddings](https://platform.openai.com/docs/guides/embeddings) via [LiteLLM](https://www.litellm.ai).
```python
import asyncio
from qdrant_client import AsyncQdrantClient
from ragbits.core.embeddings.litellm import LiteLLMEmbeddings
from ragbits.core.vector_stores.qdrant import QdrantVectorStore
from ragbits.document_search import DocumentSearch, SearchConfig
from ragbits.document_search.documents.document import DocumentMeta
documents = [
DocumentMeta.create_text_document_from_literal(
"RIP boiled water. You will be mist."
),
DocumentMeta.create_text_document_from_literal(
"Why programmers don't like to swim? Because they're scared of the floating points."
),
DocumentMeta.create_text_document_from_literal("This one is completely unrelated."),
]
async def main() -> None:
embedder = LiteLLMEmbeddings(
model="text-embedding-3-small",
)
vector_store = QdrantVectorStore(
client=AsyncQdrantClient(url="http://localhost:6333"),
collection_name="{collection_name}",
)
document_search = DocumentSearch(
embedder=embedder,
vector_store=vector_store,
)
await document_search.ingest(documents)
all_documents = await vector_store.list()
print([doc.metadata["content"] for doc in all_documents])
query = "I write computer software. Tell me something."
vector_store_kwargs = {
"k": 1,
"max_distance": None,
}
results = await document_search.search(
query,
config=SearchConfig(vector_store_kwargs=vector_store_kwargs),
)
print(f"Documents similar to: {query}")
print([element.get_key() for element in results])
```
</details>
## 📚 Further Reading
- Ragbits [Documentation](http://ragbits.deepsense.ai)
- [Source Code](https://github.com/deepsense-ai/ragbits)
@@ -1,94 +0,0 @@
---
title: Solon
---
# Solon
[Solon](https://solon.noear.org) is a lightweight, high-performance Java enterprise framework designed for efficient, eco-friendly development. It enhances concurrency, reduces memory usage, speeds up startup, minimizes packaging size, and supports Java 8 to Java 23, offering a flexible alternative to Spring.
Qdrant is available as a component in Solon-AI for efficient vector indexing and retrievals.
## Installation
```xml
<dependency>
<groupId>org.noear</groupId>
<artifactId>solon-ai-repo-qdrant</artifactId>
</dependency>
```
This is the main extension plugin for **solon-ai**, which provides the `QdrantRepository` knowledge base.
## Configuration
When using `QdrantRepository`, an embedding model needs to be configured.
```yaml
solon.ai.embed:
bgem3:
apiUrl: "http://127.0.0.1:11434/api/embed"
provider: "ollama"
model: "bge-m3:latest"
solon.ai.repo:
qdrant:
host: "localhost"
port: 6334
useSsl: false
```
You can now instantiate the embedding model and Qdrant.
```java
@Configuration
public class DemoConfig {
// Create the embedding model
@Bean
public EmbeddingModel embeddingModel(@Inject("${solon.ai.embed.bgem3}") EmbeddingConfig config) {
return EmbeddingModel.of(config).build();
}
// Configure the QdrantClient using QdrantGrpcClient
@BindProps(prefix = "solon.ai.repo.qdrant")
@Bean
public QdrantClient qdrantClient(@Value("${solon.ai.repo.qdrant.host}") String host,
@Value("${solon.ai.repo.qdrant.port}") int port,
@Value("${solon.ai.repo.qdrant.useSsl}") boolean useSsl) {
return new QdrantClient(
QdrantGrpcClient.newBuilder(host, port, useSsl).build()
);
}
// Initialize the Qdrant knowledge base
@Bean
public QdrantRepository repository(EmbeddingModel embeddingModel, QdrantClient client) {
return new QdrantRepository(embeddingModel, client);
}
}
```
## Usage
```java
@Component
public class DemoService {
@Inject
private QdrantRepository repository;
// Add documents to the repository
public void addDocument(List<Document> docs) {
repository.insert(docs);
}
// Search for documents based on a query
public List<Document> findDocument(String query) {
return repository.search(query);
}
}
```
## Next steps
- Solon [Documentation](https://solon.noear.org).
- Solon [Source](https://github.com/opensolon/solon)
@@ -1,87 +0,0 @@
---
title: Superduper
---
# Superduper
[Superduper](https://superduper.io/) is a framework for building flexible, compositional AI applications which may be applied directly to databases using a declarative programming model. These applications declare and maintain a desired state of the database, and use the database directly to store outputs of AI components, meta-data about the components and data pertaining to the state of the system.
Qdrant is available as a vector search provider in Superduper.
## Installation
```bash
pip install superduper-framework
```
## Setup
- To use Qdrant for vector search layer, create a `settings.yaml` file with the following config.
> settings.yaml
```yaml
cluster:
vector_search:
type: qdrant
vector_search_kwargs:
url: "http://localhost:6333"
api_key: "<YOUR_API_KEY>
# Supports all parameters of qdrant_client.QdrantClient
```
- Set the `SUPERDUPER_CONFIG` env value to the path of the config file.
```bash
export SUPERDUPER_CONFIG=path/to/settings.yaml
```
That's all. You can now use Superduper backed by Qdrant.
## Example
Here's an example to run vector search using the configured Qdrant index.
```python
import json
import requests
from superduper import superduper, Document
from superduper.ext.sentence_transformers import SentenceTransformer
r = requests.get('https://superduperdb-public-demo.s3.amazonaws.com/text.json')
with open('text.json', 'wb') as f:
f.write(r.content)
with open('text.json', 'r') as f:
data = json.load(f)
db = superduper('mongomock://test')
_ = db['documents'].insert_many([Document({'txt': txt}) for txt in data]).execute()
model = SentenceTransformer(
identifier="test",
predict_kwargs={"show_progress_bar": True},
model="all-MiniLM-L6-v2",
device="cpu",
postprocess=lambda x: x.tolist(),
)
vector_index = model.to_vector_index(select=db['documents'].find(), key='txt')
db.apply(vector_index)
query = db['documents'].like({'txt': 'Tell me about vector-search'}, vector_index=vector_index.identifier, n=3).find()
cursor = query.execute()
for r in cursor:
print('=' * 100)
print(r.unpack()['txt'])
print('=' * 100)
```
## 📚 Further Reading
- Superduper [Intro](https://docs.superduper.io/docs/intro)
- Superduper [Source](https://github.com/superduper-io/superduper)