Merge branch 'master' into rig-int

This commit is contained in:
David Myriel
2024-11-14 15:20:41 -08:00
committed by GitHub
128 changed files with 2431 additions and 384 deletions
@@ -1,31 +1,36 @@
---
title: Frameworks
weight: 20
partition: build
---
## Framework Integrations
| Framework | Description |
| ------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| [AutoGen](/documentation/frameworks/autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
| [Canopy](/documentation/frameworks/canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
| [Cheshire Cat](/documentation/frameworks/cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
| [DocArray](/documentation/frameworks/docarray/) | Python library for managing data in multi-modal AI applications. |
| [DSPy](/documentation/frameworks/dspy/) | Framework for algorithmically optimizing LM prompts and weights. |
| [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. |
| [Genkit](/documentation/frameworks/genkit/) | Framework to build, deploy, and monitor production-ready AI-powered apps. |
| [Haystack](/documentation/frameworks/haystack/) | LLM orchestration framework to build customizable, production-ready LLM applications. |
| [Lakechain](/documentation/frameworks/lakechain/) | Python framework for deploying document processing pipelines on AWS using infrastructure-as-code. |
| [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. |
| [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. |
| [Langchain4j](/documentation/frameworks/langchain4j/) | Java framework for building context-aware, reasoning applications using LLMs. |
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
| [MemGPT](/documentation/frameworks/memgpt/) | System to build LLM agents with long term memory & custom tools |
| [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language |
| [Rig-rs](/documentation/frameworks/rig-rs/) | Rust library for building scalable, modular, and ergonomic LLM-powered applications. |
| [Semantic Router](/documentation/frameworks/semantic-router/) | Python library to build a decision-making layer for AI applications using vector search. |
| [Spring AI](/documentation/frameworks/spring-ai/) | Java AI framework for building with Spring design principles such as portability and modular design. |
| [Sycamore](/documentation/frameworks/sycamore/) | Document processing engine for ETL, RAG, LLM-based applications, and analytics on unstructured data. |
| [txtai](/documentation/frameworks/txtai/) | Python library for semantic search, LLM orchestration and language model workflows. |
| [Vanna AI](/documentation/frameworks/vanna-ai/) | Python RAG framework for SQL generation and querying. |
| Framework | Description |
| ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| [AutoGen](/documentation/frameworks/autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
| [Canopy](/documentation/frameworks/canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
| [Cheshire Cat](/documentation/frameworks/cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
| [DocArray](/documentation/frameworks/docarray/) | Python library for managing data in multi-modal AI applications. |
| [DSPy](/documentation/frameworks/dspy/) | Framework for algorithmically optimizing LM prompts and weights. |
| [Feast](/documentation/frameworks/feast/) | Open-source feature store to operate production ML systems at scale as a set of features. |
| [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. |
| [Genkit](/documentation/frameworks/genkit/) | Framework to build, deploy, and monitor production-ready AI-powered apps. |
| [Haystack](/documentation/frameworks/haystack/) | LLM orchestration framework to build customizable, production-ready LLM applications. |
| [Lakechain](/documentation/frameworks/lakechain/) | Python framework for deploying document processing pipelines on AWS using infrastructure-as-code. |
| [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. |
| [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. |
| [Langchain4j](/documentation/frameworks/langchain4j/) | Java framework for building context-aware, reasoning applications using LLMs. |
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
| [MemGPT](/documentation/frameworks/memgpt/) | System to build LLM agents with long term memory & custom tools |
| [Neo4j GraphRAG](/documentation/frameworks/neo4j-graphrag/) | Package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. |
| [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language |
| [Ragbits](/documentation/frameworks/ragbits/) | Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. |
| [Rig-rs](/documentation/frameworks/rig-rs/) | Rust library for building scalable, modular, and ergonomic LLM-powered applications. |
| [Semantic Router](/documentation/frameworks/semantic-router/) | Python library to build a decision-making layer for AI applications using vector search. |
| [Spring AI](/documentation/frameworks/spring-ai/) | Java AI framework for building with Spring design principles such as portability and modular design. |
| [Sycamore](/documentation/frameworks/sycamore/) | Document processing engine for ETL, RAG, LLM-based applications, and analytics on unstructured data. |
| [txtai](/documentation/frameworks/txtai/) | Python library for semantic search, LLM orchestration and language model workflows. |
| [Vanna AI](/documentation/frameworks/vanna-ai/) | Python RAG framework for SQL generation and querying. |
@@ -0,0 +1,63 @@
---
title: Feast
---
## Feast
[Feast (**Fe**ature **St**ore)](https://docs.feast.dev) is an open-source feature store that helps teams operate production ML systems at scale by allowing them to define, manage, validate, and serve features for production AI/ML.
Qdrant is available as a supported vectorstore in Feast to integrate in your workflows.
## Insatallation
To use the Qdrant online store, you need to install Feast with the `qdrant` extra.
```bash
pip install 'feast[qdrant]'
```
## Usage
An example config with Qdrant could look like:
```yaml
project: my_feature_repo
registry: data/registry.db
provider: local
online_store:
type: qdrant
host: xyz-example.eu-central.aws.cloud.qdrant.io
port: 6333
api_key: <your-own-key>
vector_len: 384
# Reference: https://qdrant.tech/documentation/concepts/vectors/#named-vectors
# vector_name: text-vec
write_batch_size: 100
```
You can refer to the Feast [reference](https://rtd.feast.dev/en/master/index.html#) for the full list of configuration options.
## Retrieving Documents
The Qdrant online store supports retrieving document vectors for a given list of entity keys. The document vectors are returned as a dictionary where the key is the entity key and the value being the vector.
```python
from feast import FeatureStore
feature_store = FeatureStore(repo_path="feature_store.yaml")
query_vector = [1.0, 2.0, 3.0, 4.0, 5.0]
top_k = 5
feature_values = feature_store.retrieve_online_documents(
feature="my_feature",
query=query_vector,
top_k=top_k
)
```
## 📚 Further Reading
- [Feast Docs](http://docs.feast.dev/)
- [Feast Reference](https://rtd.feast.dev/en/master/index.html/)
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/)
@@ -0,0 +1,69 @@
---
title: Neo4j GraphRAG
---
# Neo4j GraphRAG
[Neo4j GraphRAG](https://neo4j.com/docs/neo4j-graphrag-python/current/) is a Python package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. As a first-party library, it offers a robust, feature-rich, and high-performance solution, with the added assurance of long-term support and maintenance directly from Neo4j. It offers a Qdrant retriever natively to search for vectors stored in a Qdrant collection.
## Installation
```bash
pip install neo4j-graphrag[qdrant]
```
## Usage
A vector query with Neo4j and Qdrant could look like:
```python
from neo4j import GraphDatabase
from neo4j_graphrag.retrievers import QdrantNeo4jRetriever
from qdrant_client import QdrantClient
from examples.embedding_biology import EMBEDDING_BIOLOGY
NEO4J_URL = "neo4j://localhost:7687"
NEO4J_AUTH = ("neo4j", "password")
with GraphDatabase.driver(NEO4J_URL, auth=NEO4J_AUTH) as neo4j_driver:
retriever = QdrantNeo4jRetriever(
driver=neo4j_driver,
client=QdrantClient(url="http://localhost:6333"),
collection_name="{collection_name}",
id_property_external="neo4j_id",
id_property_neo4j="id",
)
retriever.search(query_vector=[0.5523, 0.523, 0.132, 0.523, ...], top_k=5)
```
Alternatively, you can use any [Langchain embeddings providers](https://python.langchain.com/docs/integrations/text_embedding/), to vectorize text queries automatically.
```python
from langchain_huggingface.embeddings import HuggingFaceEmbeddings
from neo4j import GraphDatabase
from neo4j_graphrag.retrievers import QdrantNeo4jRetriever
from qdrant_client import QdrantClient
NEO4J_URL = "neo4j://localhost:7687"
NEO4J_AUTH = ("neo4j", "password")
with GraphDatabase.driver(NEO4J_URL, auth=NEO4J_AUTH) as neo4j_driver:
embedder = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
retriever = QdrantNeo4jRetriever(
driver=neo4j_driver,
client=QdrantClient(url="http://localhost:6333"),
collection_name="{collection_name}",
id_property_external="neo4j_id",
id_property_neo4j="id",
embedder=embedder,
)
retriever.search(query_text="my user query", top_k=10)
```
## Further Reading
- [Neo4j GraphRAG Reference](https://neo4j.com/docs/neo4j-graphrag-python/current/index.html)
- [Qdrant Retriever Reference](https://neo4j.com/docs/neo4j-graphrag-python/current/user_guide_rag.html#qdrant-neo4j-retriever-user-guide)
- [Source](https://github.com/neo4j/neo4j-graphrag-python/tree/main/src/neo4j_graphrag/retrievers/external/qdrant)
@@ -0,0 +1,83 @@
---
title: Ragbits
---
# Ragbits
[Ragbit](https://ragbits.deepsense.ai) is a Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. It prioritizes developer experience by providing a simple and intuitive API. It also includes a comprehensive set of tools for seamlessly building, testing, and deploying your RAG applications efficiently.
Qdrant is available as a vectorstore in Ragbits to ingest and search search documents from a collection.
## Installation
Install the Python package that comes bundled with the Qdrant integration.
```bash
pip install ragbits
```
## Usage
An example usage of Ragbits and Qdrant would look something like this:
The following example uses [OpenAI embeddings](https://platform.openai.com/docs/guides/embeddings) via [LiteLLM](https://www.litellm.ai).
```python
import asyncio
from qdrant_client import AsyncQdrantClient
from ragbits.core.embeddings.litellm import LiteLLMEmbeddings
from ragbits.core.vector_stores.qdrant import QdrantVectorStore
from ragbits.document_search import DocumentSearch, SearchConfig
from ragbits.document_search.documents.document import DocumentMeta
documents = [
DocumentMeta.create_text_document_from_literal(
"RIP boiled water. You will be mist."
),
DocumentMeta.create_text_document_from_literal(
"Why programmers don't like to swim? Because they're scared of the floating points."
),
DocumentMeta.create_text_document_from_literal("This one is completely unrelated."),
]
async def main() -> None:
embedder = LiteLLMEmbeddings(
model="text-embedding-3-small",
)
vector_store = QdrantVectorStore(
client=AsyncQdrantClient(url="http://localhost:6333"),
collection_name="{collection_name}",
)
document_search = DocumentSearch(
embedder=embedder,
vector_store=vector_store,
)
await document_search.ingest(documents)
all_documents = await vector_store.list()
print([doc.metadata["content"] for doc in all_documents])
query = "I write computer software. Tell me something."
vector_store_kwargs = {
"k": 1,
"max_distance": None,
}
results = await document_search.search(
query,
config=SearchConfig(vector_store_kwargs=vector_store_kwargs),
)
print(f"Documents similar to: {query}")
print([element.get_key() for element in results])
```
</details>
## 📚 Further Reading
- Ragbits [Documentation](http://ragbits.deepsense.ai)
- [Source Code](https://github.com/deepsense-ai/ragbits)