mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-29 07:58:31 +02:00
Add Gemini Embedding Model 001 (#457)
* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant * * refactor(integrations): move cohere.md to embeddings folder * refactor(integrations): move openai.md to embeddings folder * refactor(integrations): move autogen.md to frameworks folder * refactor(integrations): move langchain.md to frameworks folder * blacken * * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both * * chore(gemini.md): remove old Gemini integration documentation * * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false * chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false * Split integrations into embedding and frameworks * Update heading level for embedding a document * Update Gemini embedding documentation * Update titles for embedding and frameworks sections * Try again with nesting * Add documentation for integrated frameworks and embedding options * Delete integrations documentation file * Add Delimiter; unknown weights * Change all weights to 3x * Delimiter reorg * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * * docs(embedding/gemini.md): update Gemini Embedding Model API documentation * * - Add information about the new Gemini Embedding Model and its compatibility with Qdrant * - Clarify the usage of the `task_type` parameter in the API call * - Provide a list of supported task types and * * docs(embedding): update list of embedding integrations * * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md * refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md * * chore(embedding): update is_empty value to true in _index.md * chore(embedding): remove Fifty One from embedding/_index.md * embedding -> embeddings --------- Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
This commit is contained in:
@@ -0,0 +1,24 @@
|
||||
---
|
||||
title: Frameworks
|
||||
weight: 33
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: true
|
||||
---
|
||||
|
||||
| Frameworks |
|
||||
|---|
|
||||
| [AirByte](./airbyte/) |
|
||||
| [AutoGen](./autogen/) |
|
||||
| [Cheshire Cat](./cheshire-cat/) |
|
||||
| [DLT](./dlt/) |
|
||||
| [DocArray](./docarray/) |
|
||||
| [DSPy](./dspy/) |
|
||||
| [Fifty One](./fifty-one/) |
|
||||
| [txtai](./txtai/) |
|
||||
| [Fondant](./fondant/) |
|
||||
| [Haystack](./haystack/) |
|
||||
| [Langchain](./langchain/) |
|
||||
| [Llama Index](./llama-index/) |
|
||||
| [Minds DB](./mindsdb/) |
|
||||
| [PrivateGPT](./privategpt/) |
|
||||
| [Spark](./spark/) |
|
||||
@@ -0,0 +1,77 @@
|
||||
---
|
||||
title: Airbyte
|
||||
weight: 1000
|
||||
---
|
||||
|
||||
# Airbyte
|
||||
|
||||
[Airbyte](https://airbyte.com/) is an open-source data integration platform that helps you replicate your data
|
||||
between different systems. It has a [growing list of connectors](https://docs.airbyte.io/integrations) that can
|
||||
be used to ingest data from multiple sources. Building data pipelines is also crucial for managing the data in
|
||||
Qdrant, and Airbyte is a great tool for this purpose.
|
||||
|
||||
Airbyte may take care of the data ingestion from a selected source, while Qdrant will help you to build a search
|
||||
engine on top of it. There are three supported modes of how the data can be ingested into Qdrant:
|
||||
|
||||
* **Full Refresh Sync**
|
||||
* **Incremental - Append Sync**
|
||||
* **Incremental - Append + Deduped**
|
||||
|
||||
You can read more about these modes in the [Airbyte documentation](https://docs.airbyte.io/integrations/destinations/qdrant).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before you start, make sure you have the following:
|
||||
|
||||
1. Airbyte instance, either [Open Source](https://airbyte.com/solutions/airbyte-open-source),
|
||||
[Self-Managed](https://airbyte.com/solutions/airbyte-enterprise), or [Cloud](https://airbyte.com/solutions/airbyte-cloud).
|
||||
2. Running instance of Qdrant. It has to be accessible by URL from the machine where Airbyte is running.
|
||||
You can follow the [installation guide](/documentation/guides/installation/) to set up Qdrant.
|
||||
|
||||
## Setting up Qdrant as a destination
|
||||
|
||||
Once you have a running instance of Airbyte, you can set up Qdrant as a destination directly in the UI.
|
||||
Airbyte's Qdrant destination is connected with a single collection in Qdrant.
|
||||
|
||||

|
||||
|
||||
### Text processing
|
||||
|
||||
Airbyte has some built-in mechanisms to transform your texts into embeddings. You can choose how you want to
|
||||
chunk your fields into pieces before calculating the embeddings, but also which fields should be used to
|
||||
create the point payload.
|
||||
|
||||

|
||||
|
||||
### Embeddings
|
||||
|
||||
You can choose the model that will be used to calculate the embeddings. Currently, Airbyte supports multiple
|
||||
models, including OpenAI and Cohere.
|
||||
|
||||

|
||||
|
||||
Using some precomputed embeddings from your data source is also possible. In this case, you can pass the field
|
||||
name containing the embeddings and their dimensionality.
|
||||
|
||||

|
||||
|
||||
### Qdrant connection details
|
||||
|
||||
Finally, we can configure the target Qdrant instance and collection. In case you use the built-in authentication
|
||||
mechanism, here is where you can pass the token.
|
||||
|
||||

|
||||
|
||||
Once you confirm creating the destination, Airbyte will test if a specified Qdrant cluster is accessible and
|
||||
might be used as a destination.
|
||||
|
||||
## Setting up connection
|
||||
|
||||
Airbyte combines sources and destinations into a single entity called a connection. Once you have a destination
|
||||
configured and a source, you can create a connection between them. It doesn't matter what source you use, as
|
||||
long as Airbyte supports it. The process is pretty straightforward, but depends on the source you use.
|
||||
|
||||

|
||||
|
||||
More information about creating connections can be found in the
|
||||
[Airbyte documentation](https://docs.airbyte.com/understanding-airbyte/connections/).
|
||||
@@ -0,0 +1,102 @@
|
||||
---
|
||||
title: Autogen
|
||||
weight: 1200
|
||||
---
|
||||
|
||||
# Microsoft Autogen
|
||||
|
||||
[AutoGen](https://github.com/microsoft/autogen) is a framework that enables the development of LLM applications using multiple agents that can converse with each other to solve tasks. AutoGen agents are customizable, conversable, and seamlessly allow human participation. They can operate in various modes that employ combinations of LLMs, human inputs, and tools.
|
||||
|
||||
- Multi-agent conversations: AutoGen agents can communicate with each other to solve tasks. This allows for more complex and sophisticated applications than would be possible with a single LLM.
|
||||
- Customization: AutoGen agents can be customized to meet the specific needs of an application. This includes the ability to choose the LLMs to use, the types of human input to allow, and the tools to employ.
|
||||
- Human participation: AutoGen seamlessly allows human participation. This means that humans can provide input and feedback to the agents as needed.
|
||||
|
||||
With the Autogen-Qdrant integration, you can use the `QdrantRetrieveUserProxyAgent` from autogen to build retrieval augmented generation(RAG) services with ease.
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
pip install "pyautogen[retrievechat]" "qdrant_client[fastembed]"
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
A demo application that generates code based on context w/o human feedback
|
||||
|
||||
#### Set your API Endpoint
|
||||
|
||||
The config_list_from_json function loads a list of configurations from an environment variable or a JSON file.
|
||||
|
||||
```python
|
||||
from autogen import config_list_from_json
|
||||
from autogen.agentchat.contrib.retrieve_assistant_agent import RetrieveAssistantAgent
|
||||
from autogen.agentchat.contrib.qdrant_retrieve_user_proxy_agent import QdrantRetrieveUserProxyAgent
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
config_list = config_list_from_json(
|
||||
env_or_file="OAI_CONFIG_LIST",
|
||||
file_location="."
|
||||
)
|
||||
```
|
||||
|
||||
It first looks for the environment variable "OAI_CONFIG_LIST" which needs to be a valid JSON string. If that variable is not found, it then looks for a JSON file named "OAI_CONFIG_LIST". The file structure sample can be found [here](https://github.com/microsoft/autogen/blob/main/OAI_CONFIG_LIST_sample).
|
||||
|
||||
#### Construct agents for RetrieveChat
|
||||
|
||||
We start by initializing the RetrieveAssistantAgent and QdrantRetrieveUserProxyAgent. The system message needs to be set to "You are a helpful assistant." for RetrieveAssistantAgent. The detailed instructions are given in the user message.
|
||||
|
||||
```python
|
||||
# Print the generation steps
|
||||
autogen.ChatCompletion.start_logging()
|
||||
|
||||
# 1. create a RetrieveAssistantAgent instance named "assistant"
|
||||
assistant = RetrieveAssistantAgent(
|
||||
name="assistant",
|
||||
system_message="You are a helpful assistant.",
|
||||
llm_config={
|
||||
"request_timeout": 600,
|
||||
"seed": 42,
|
||||
"config_list": config_list,
|
||||
},
|
||||
)
|
||||
|
||||
# 2. create a QdrantRetrieveUserProxyAgent instance named "qdrantagent"
|
||||
# By default, the human_input_mode is "ALWAYS", i.e. the agent will ask for human input at every step.
|
||||
# `docs_path` is the path to the docs directory.
|
||||
# `task` indicates the kind of task we're working on.
|
||||
# `chunk_token_size` is the chunk token size for the retrieve chat.
|
||||
# We use an in-memory QdrantClient instance here. Not recommended for production.
|
||||
|
||||
ragproxyagent = QdrantRetrieveUserProxyAgent(
|
||||
name="qdrantagent",
|
||||
human_input_mode="NEVER",
|
||||
max_consecutive_auto_reply=10,
|
||||
retrieve_config={
|
||||
"task": "code",
|
||||
"docs_path": "./path/to/docs",
|
||||
"chunk_token_size": 2000,
|
||||
"model": config_list[0]["model"],
|
||||
"client": QdrantClient(":memory:"),
|
||||
"embedding_model": "BAAI/bge-small-en-v1.5",
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
#### Run the retriever service
|
||||
|
||||
```python
|
||||
# Always reset the assistant before starting a new conversation.
|
||||
assistant.reset()
|
||||
|
||||
# We use the ragproxyagent to generate a prompt to be sent to the assistant as the initial message.
|
||||
# The assistant receives the message and generates a response. The response will be sent back to the ragproxyagent for processing.
|
||||
# The conversation continues until the termination condition is met, in RetrieveChat, the termination condition when no human-in-loop is no code block detected.
|
||||
|
||||
# The query used below is for demonstration. It should usually be related to the docs made available to the agent
|
||||
code_problem = "How can I use FLAML to perform a classification task?"
|
||||
ragproxyagent.initiate_chat(assistant, problem=code_problem)
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
Check out more Autogen [examples](https://microsoft.github.io/autogen/docs/Examples/AgentChat). You can find detailed documentation about AutoGen [here](https://microsoft.github.io/autogen/).
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
title: Cheshire Cat
|
||||
weight: 600
|
||||
---
|
||||
|
||||
# Cheshire Cat
|
||||
|
||||
[Cheshire Cat](https://cheshirecat.ai/) is an open-source framework that allows you to develop intelligent agents on top of many Large Language Models (LLM). You can develop your custom AI architecture to assist you in a wide range of tasks.
|
||||
|
||||

|
||||
|
||||
## Cheshire Cat and Qdrant
|
||||
|
||||
Cheshire Cat uses Qdrant as the default [Vector Memory](https://cheshire-cat-ai.github.io/docs/conceptual/memory/vector_memory/) for ingesting and retrieving documents.
|
||||
|
||||
```
|
||||
# Decide host and port for your Cat. Default will be localhost:1865
|
||||
CORE_HOST=localhost
|
||||
CORE_PORT=1865
|
||||
|
||||
# Qdrant server
|
||||
# QDRANT_HOST=localhost
|
||||
# QDRANT_PORT=6333
|
||||
```
|
||||
|
||||
Cheshire Cat takes great advantage of the following features of Qdrant:
|
||||
* [Collection Aliases](../../concepts/collections/#collection-aliases) to manage the change from one embedder to another.
|
||||
* [Quantization](../../guides/quantization/) to obtain a good balance between speed, memory usage and quality of the results.
|
||||
* [Snapshots](../../concepts/snapshots/) to not miss any information.
|
||||
* [Community](https://discord.com/invite/tdtYvXjC4h)
|
||||
|
||||

|
||||
|
||||
## How to use the Cheshire Cat
|
||||
|
||||
### Requirements
|
||||
To run the Cheshire Cat, you need to have [Docker](https://docs.docker.com/engine/install/) and [docker-compose](https://docs.docker.com/compose/install/) already installed on your system.
|
||||
|
||||
```shell
|
||||
docker run --rm -it -p 1865:80 ghcr.io/cheshire-cat-ai/core:latest
|
||||
```
|
||||
|
||||
* Chat with the Cheshire Cat on [localhost:1865/admin](http://localhost:1865/admin).
|
||||
* You can also interact via REST API and try out the endpoints on [localhost:1865/docs](http://localhost:1865/docs)
|
||||
|
||||
Check the [instructions on github](https://github.com/cheshire-cat-ai/core/blob/main/README.md) for a more comprehensive quick start.
|
||||
|
||||
### First configuration of the LLM
|
||||
|
||||
* Open the Admin Portal in your browser at [localhost:1865/admin](http://localhost:1865/admin).
|
||||
* Configure the LLM in the `Settings` tab.
|
||||
* If you don't explicitly choose it using `Settings` tab, the Embedder follows the LLM.
|
||||
|
||||
## Next steps
|
||||
|
||||
For more information, refer to the Cheshire Cat [documentation](https://cheshire-cat-ai.github.io/docs/) and [blog](https://cheshirecat.ai/blog/).
|
||||
|
||||
* [Getting started](https://cheshirecat.ai/hello-world/)
|
||||
* [How the Cat works](https://cheshirecat.ai/how-the-cat-works/)
|
||||
* [Write Your First Plugin](https://cheshirecat.ai/write-your-first-plugin/)
|
||||
* [Cheshire Cat's use of Qdrant - Vector Space](https://cheshirecat.ai/dont-get-lost-in-vector-space/)
|
||||
* [Cheshire Cat's use of Qdrant - Aliases](https://cheshirecat.ai/the-drunken-cat-effect/)
|
||||
* [Discord Community](https://discord.com/invite/bHX5sNFCYU)
|
||||
@@ -0,0 +1,101 @@
|
||||
---
|
||||
title: DLT
|
||||
weight: 1300
|
||||
---
|
||||
|
||||
# DLT(Data Load Tool)
|
||||
|
||||
[DLT](https://dlthub.com/) is an open-source library that you can add to your Python scripts to load data from various and often messy data sources into well-structured, live datasets.
|
||||
|
||||
With the DLT-Qdrant integration, you can now select Qdrant as a DLT destination to load data into.
|
||||
|
||||
**DLT Enables**
|
||||
|
||||
- Automated maintenance - with schema inference, alerts and short declarative code, maintenance becomes simple.
|
||||
- Run it where Python runs - on Airflow, serverless functions, notebooks. Scales on micro and large infrastructure alike.
|
||||
- User-friendly, declarative interface that removes knowledge obstacles for beginners while empowering senior professionals.
|
||||
|
||||
## Usage
|
||||
|
||||
To get started, install `dlt` with the `qdrant` extra.
|
||||
|
||||
```bash
|
||||
pip install "dlt[qdrant]"
|
||||
```
|
||||
|
||||
Configure the destination in the DLT secrets file. The file is located at `~/.dlt/secrets.toml` by default. Add the following section to the secrets file.
|
||||
|
||||
```toml
|
||||
[destination.qdrant.credentials]
|
||||
location = "https://your-qdrant-url"
|
||||
api_key = "your-qdrant-api-key"
|
||||
```
|
||||
|
||||
The location will default to `http://localhost:6333` and `api_key` is not defined - which are the defaults for a local Qdrant instance.
|
||||
Find more information about DLT configurations [here](https://dlthub.com/docs/general-usage/credentials).
|
||||
|
||||
Define the source of the data.
|
||||
|
||||
```python
|
||||
import dlt
|
||||
from dlt.destinations.qdrant import qdrant_adapter
|
||||
|
||||
movies = [
|
||||
{
|
||||
"title": "Blade Runner",
|
||||
"year": 1982,
|
||||
"description": "The film is about a dystopian vision of the future that combines noir elements with sci-fi imagery."
|
||||
},
|
||||
{
|
||||
"title": "Ghost in the Shell",
|
||||
"year": 1995,
|
||||
"description": "The film is about a cyborg policewoman and her partner who set out to find the main culprit behind brain hacking, the Puppet Master."
|
||||
},
|
||||
{
|
||||
"title": "The Matrix",
|
||||
"year": 1999,
|
||||
"description": "The movie is set in the 22nd century and tells the story of a computer hacker who joins an underground group fighting the powerful computers that rule the earth."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
<aside role="status">
|
||||
A more comprehensive pipeline would load data from some API or use one of <a href="https://dlthub.com/docs/dlt-ecosystem/verified-sources">DLT's verified sources</a>.
|
||||
</aside>
|
||||
|
||||
Define the pipeline.
|
||||
|
||||
```python
|
||||
pipeline = dlt.pipeline(
|
||||
pipeline_name="movies",
|
||||
destination="qdrant",
|
||||
dataset_name="movies_dataset",
|
||||
)
|
||||
```
|
||||
|
||||
Run the pipeline.
|
||||
|
||||
```python
|
||||
info = pipeline.run(
|
||||
qdrant_adapter(
|
||||
movies,
|
||||
embed=["title", "description"]
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
The data is now loaded into Qdrant.
|
||||
|
||||
To use vector search after the data has been loaded, you must specify which fields Qdrant needs to generate embeddings for. You do that by wrapping the data (or [DLT resource](https://dlthub.com/docs/general-usage/resource)) with the `qdrant_adapter` function.
|
||||
|
||||
## Write disposition
|
||||
|
||||
A DLT [write disposition](https://dlthub.com/docs/dlt-ecosystem/destinations/qdrant/#write-disposition) defines how the data should be written to the destination. All write dispositions are supported by the Qdrant destination.
|
||||
|
||||
## DLT Sync
|
||||
|
||||
Qdrant destination supports syncing of the [`DLT` state](https://dlthub.com/docs/general-usage/state#syncing-state-with-destination).
|
||||
|
||||
## Next steps
|
||||
|
||||
- The comprehensive Qdrant DLT destination documentation can be found [here](https://dlthub.com/docs/dlt-ecosystem/destinations/qdrant/).
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
title: DocArray
|
||||
weight: 300
|
||||
---
|
||||
|
||||
# DocArray
|
||||
You can use Qdrant natively in DocArray, where Qdrant serves as a high-performance document store to enable scalable vector search.
|
||||
|
||||
DocArray is a library from Jina AI for nested, unstructured data in transit, including text, image, audio, video, 3D mesh, etc.
|
||||
It allows deep-learning engineers to efficiently process, embed, search, recommend, store, and transfer the data with a Pythonic API.
|
||||
|
||||
|
||||
To install DocArray with Qdrant support, please do
|
||||
|
||||
```bash
|
||||
pip install "docarray[qdrant]"
|
||||
```
|
||||
|
||||
More information can be found in [DocArray's documentations](https://docarray.jina.ai/advanced/document-store/qdrant/).
|
||||
@@ -0,0 +1,70 @@
|
||||
---
|
||||
title: Stanford DSPy
|
||||
weight: 1500
|
||||
---
|
||||
|
||||
# Stanford DSPy
|
||||
|
||||
[DSPy](https://github.com/stanfordnlp/dspy) is the framework for solving advanced tasks with language models (LMs) and retrieval models (RMs). It unifies techniques for prompting and fine-tuning LMs — and approaches for reasoning, self-improvement, and augmentation with retrieval and tools.
|
||||
|
||||
- Provides composable and declarative modules for instructing LMs in a familiar Pythonic syntax.
|
||||
|
||||
- Introduces an automatic compiler that teaches LMs how to conduct the declarative steps in your program.
|
||||
|
||||
Qdrant can be used as a retrieval mechanism in the DSPy flow.
|
||||
|
||||
## Installation
|
||||
|
||||
For the Qdrant retrieval integration, include `dspy-ai` with the `qdrant` extra:
|
||||
```bash
|
||||
pip install dspy-ai[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
We can configure `DSPy` settings to use the Qdrant retriever model like so:
|
||||
```python
|
||||
import dspy
|
||||
from dspy.retrieve.qdrant_rm import QdrantRM
|
||||
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
turbo = dspy.OpenAI(model="gpt-3.5-turbo")
|
||||
qdrant_client = QdrantClient() # Defaults to a local instance at http://localhost:6333/
|
||||
qdrant_retriever_model = QdrantRM("collection-name", qdrant_client, k=3)
|
||||
|
||||
dspy.settings.configure(lm=turbo, rm=qdrant_retriever_model)
|
||||
```
|
||||
Using the retriever is pretty simple. The `dspy.Retrieve(k)` module will search for the top-k passages that match a given query.
|
||||
|
||||
```python
|
||||
retrieve = dspy.Retrieve(k=3)
|
||||
question = "Some question about my data"
|
||||
topK_passages = retrieve(question).passages
|
||||
|
||||
print(f"Top {retrieve.k} passages for question: {question} \n", "\n")
|
||||
|
||||
for idx, passage in enumerate(topK_passages):
|
||||
print(f"{idx+1}]", passage, "\n")
|
||||
```
|
||||
|
||||
With Qdrant configured as the retriever for contexts, you can set up a DSPy module like so:
|
||||
```python
|
||||
class RAG(dspy.Module):
|
||||
def __init__(self, num_passages=3):
|
||||
super().__init__()
|
||||
|
||||
self.retrieve = dspy.Retrieve(k=num_passages)
|
||||
...
|
||||
|
||||
def forward(self, question):
|
||||
context = self.retrieve(question).passages
|
||||
...
|
||||
|
||||
```
|
||||
|
||||
With the generic RAG blueprint now in place, you can add the many interactions offered by DSPy with context retrieval powered by Qdrant.
|
||||
|
||||
## Next steps
|
||||
|
||||
Find DSPy usage docs and examples [here](https://github.com/stanfordnlp/dspy#4-documentation--tutorials).
|
||||
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: FiftyOne
|
||||
weight: 600
|
||||
---
|
||||
|
||||
# FiftyOne
|
||||
|
||||
[FiftyOne](https://voxel51.com/) is an open-source toolkit designed to enhance computer vision workflows by optimizing dataset quality
|
||||
and providing valuable insights about your models. FiftyOne 0.20, which includes a native integration with Qdrant, supporting workflows
|
||||
like [image similarity search](https://docs.voxel51.com/user_guide/brain.html#image-similarity) and
|
||||
[text search](https://docs.voxel51.com/user_guide/brain.html#text-similarity).
|
||||
|
||||
Qdrant helps FiftyOne to find the most similar images in the dataset using vector embeddings.
|
||||
|
||||
FiftyOne is available as a Python package that might be installed in the following way:
|
||||
|
||||
```bash
|
||||
pip install fiftyone
|
||||
```
|
||||
|
||||
Please check out the documentation of FiftyOne on [Qdrant integration](https://docs.voxel51.com/integrations/qdrant.html).
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
---
|
||||
title: ML6 Fondant
|
||||
weight: 1700
|
||||
---
|
||||
|
||||
# ML6 Fondant
|
||||
|
||||
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed up large-scale data processing by making containerized components reusable across pipelines and execution environments.
|
||||
|
||||
Fondant features a Qdrant component image to load textual data and embeddings into a the database.
|
||||
|
||||
## Usage
|
||||
|
||||
<aside role="status">
|
||||
A Qdrant collection has to be <a href="documentation/concepts/collections/">created in advance</a>
|
||||
</aside>
|
||||
|
||||
**A data load pipeline for RAG using Qdrant**.
|
||||
|
||||
```python
|
||||
from fondant.pipeline import ComponentOp, Pipeline
|
||||
|
||||
pipeline = Pipeline(
|
||||
pipeline_name="ingestion-pipeline",
|
||||
pipeline_description="Pipeline to prepare and process \
|
||||
data for building a RAG solution",
|
||||
base_path="./data-dir",
|
||||
)
|
||||
|
||||
# An example data source component
|
||||
load_from_source = ComponentOp(
|
||||
component_dir="path/to/data-source-component",
|
||||
arguments={
|
||||
"n_rows_to_load": 10,
|
||||
# Custom arguments for the component
|
||||
},
|
||||
)
|
||||
|
||||
chunk_text_op = ComponentOp.from_registry(
|
||||
name="chunk_text",
|
||||
arguments={
|
||||
"chunk_size": 512,
|
||||
"chunk_overlap": 32,
|
||||
},
|
||||
)
|
||||
|
||||
embed_text_op = ComponentOp.from_registry(
|
||||
name="embed_text",
|
||||
arguments={
|
||||
"model_provider": "huggingface",
|
||||
"model": "all-MiniLM-L6-v2",
|
||||
},
|
||||
)
|
||||
|
||||
# Getting the Qdrant component from the Fondant registry
|
||||
index_qdrant_op = ComponentOp.from_registry(
|
||||
name="index_qdrant",
|
||||
arguments={
|
||||
"url": "http:localhost:6333",
|
||||
"collection_name": "some-collection-name",
|
||||
},
|
||||
)
|
||||
|
||||
# Construct your pipeline
|
||||
pipeline.add_op(load_from_source)
|
||||
pipeline.add_op(chunk_text_op, dependencies=load_from_source)
|
||||
pipeline.add_op(embed_text_op, dependencies=chunk_text_op)
|
||||
pipeline.add_op(index_qdrant_op, dependencies=embed_text_op)
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
Find the FondantAI docs [here](https://fondant.ai/en/stable/).
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
title: Haystack
|
||||
weight: 400
|
||||
---
|
||||
|
||||
# Haystack
|
||||
|
||||
[Haystack](https://haystack.deepset.ai/) serves as a comprehensive NLP framework, offering a modular methodology for constructing
|
||||
cutting-edge generative AI, QA, and semantic knowledge base search systems. A critical element in contemporary NLP systems is an
|
||||
efficient database for storing and retrieving extensive text data. Vector databases excel in this role, as they house vector
|
||||
representations of text and implement effective methods for swift retrieval. Thus, we are happy to announce the integration
|
||||
with Haystack - `QdrantDocumentStore`. This document store is unique, as it is maintained externally by the Qdrant team.
|
||||
|
||||
The new document store comes as a separate package and can be updated independently of Haystack:
|
||||
|
||||
```bash
|
||||
pip install qdrant-haystack
|
||||
```
|
||||
|
||||
`QdrantDocumentStore` supports [all the configuration properties](/documentation/collections/#create-collection) available in
|
||||
the Qdrant Python client. If you want to customize the default configuration of the collection used under the hood, you can
|
||||
provide that settings when you create an instance of the `QdrantDocumentStore`. For example, if you'd like to enable the
|
||||
Scalar Quantization, you'd make that in the following way:
|
||||
|
||||
```python
|
||||
from qdrant_haystack.document_stores import QdrantDocumentStore
|
||||
from qdrant_client.http import models
|
||||
|
||||
document_store = QdrantDocumentStore(
|
||||
":memory:",
|
||||
index="Document",
|
||||
embedding_dim=512,
|
||||
recreate_index=True,
|
||||
quantization_config=models.ScalarQuantization(
|
||||
scalar=models.ScalarQuantizationConfig(
|
||||
type=models.ScalarType.INT8,
|
||||
quantile=0.99,
|
||||
always_ram=True,
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
@@ -0,0 +1,107 @@
|
||||
---
|
||||
title: LangChain
|
||||
weight: 100
|
||||
---
|
||||
|
||||
# LangChain
|
||||
|
||||
LangChain is a library that makes developing Large Language Models based applications much easier. It unifies the interfaces
|
||||
to different libraries, including major embedding providers and Qdrant. Using LangChain, you can focus on the business value
|
||||
instead of writing the boilerplate.
|
||||
|
||||
Langchain comes with the Qdrant integration by default. It might be installed with pip:
|
||||
|
||||
```bash
|
||||
pip install langchain
|
||||
```
|
||||
|
||||
Qdrant acts as a vector index that may store the embeddings with the documents used to generate them. There are various ways
|
||||
how to use it, but calling `Qdrant.from_texts` is probably the most straightforward way how to get started:
|
||||
|
||||
```python
|
||||
from langchain.vectorstores import Qdrant
|
||||
from langchain.embeddings import HuggingFaceEmbeddings
|
||||
|
||||
embeddings = HuggingFaceEmbeddings(
|
||||
model_name="sentence-transformers/all-mpnet-base-v2"
|
||||
)
|
||||
doc_store = Qdrant.from_texts(
|
||||
texts, embeddings, url="<qdrant-url>", api_key="<qdrant-api-key>", collection_name="texts"
|
||||
)
|
||||
```
|
||||
|
||||
Calling `Qdrant.from_documents` or `Qdrant.from_texts` will always recreate the collection and remove all the existing points.
|
||||
That's fine for some experiments, but you'll prefer not to start from scratch every single time in a real-world scenario.
|
||||
If you prefer reusing an existing collection, you can create an instance of Qdrant on your own:
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
|
||||
embeddings = HuggingFaceEmbeddings(
|
||||
model_name="sentence-transformers/all-mpnet-base-v2"
|
||||
)
|
||||
|
||||
client = qdrant_client.QdrantClient(
|
||||
"<qdrant-url>",
|
||||
api_key="<qdrant-api-key>", # For Qdrant Cloud, None for local instance
|
||||
)
|
||||
|
||||
doc_store = Qdrant(
|
||||
client=client, collection_name="texts",
|
||||
embeddings=embeddings,
|
||||
)
|
||||
```
|
||||
|
||||
## Local mode
|
||||
|
||||
Python client allows you to run the same code in local mode without running the Qdrant server. That's great for testing things
|
||||
out and debugging or if you plan to store just a small amount of vectors. The embeddings might be fully kepy in memory or
|
||||
persisted on disk.
|
||||
|
||||
### In-memory
|
||||
|
||||
For some testing scenarios and quick experiments, you may prefer to keep all the data in memory only, so it gets lost when the
|
||||
client is destroyed - usually at the end of your script/notebook.
|
||||
|
||||
```python
|
||||
qdrant = Qdrant.from_documents(
|
||||
docs, embeddings,
|
||||
location=":memory:", # Local mode with in-memory storage only
|
||||
collection_name="my_documents",
|
||||
)
|
||||
```
|
||||
|
||||
### On-disk storage
|
||||
|
||||
Local mode, without using the Qdrant server, may also store your vectors on disk so they’re persisted between runs.
|
||||
|
||||
```python
|
||||
qdrant = Qdrant.from_documents(
|
||||
docs, embeddings,
|
||||
path="/tmp/local_qdrant",
|
||||
collection_name="my_documents",
|
||||
)
|
||||
```
|
||||
|
||||
### On-premise server deployment
|
||||
|
||||
No matter if you choose to launch Qdrant locally with [a Docker container](/documentation/guides/installation/), or
|
||||
select a Kubernetes deployment with [the official Helm chart](https://github.com/qdrant/qdrant-helm), the way you're
|
||||
going to connect to such an instance will be identical. You'll need to provide a URL pointing to the service.
|
||||
|
||||
```python
|
||||
url = "<---qdrant url here --->"
|
||||
qdrant = Qdrant.from_documents(
|
||||
docs,
|
||||
embeddings,
|
||||
url,
|
||||
prefer_grpc=True,
|
||||
collection_name="my_documents",
|
||||
)
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
If you'd like to know more about running Qdrant in a LangChain-based application, please read our article
|
||||
[Question Answering with LangChain and Qdrant without boilerplate](/articles/langchain-integration/). Some more information
|
||||
might also be found in the [LangChain documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant).
|
||||
@@ -0,0 +1,34 @@
|
||||
---
|
||||
title: LlamaIndex
|
||||
weight: 200
|
||||
---
|
||||
|
||||
# LlamaIndex (GPT Index)
|
||||
|
||||
LlamaIndex (formerly GPT Index) acts as an interface between your external data and Large Language Models. So you can bring your
|
||||
private data and augment LLMs with it. LlamaIndex simplifies data ingestion and indexing, integrating Qdrant as a vector index.
|
||||
|
||||
Installing LlamaIndex is straightforward if we use pip as a package manager. Qdrant is not installed by default, so we need to
|
||||
install it separately:
|
||||
|
||||
```bash
|
||||
pip install llama-index qdrant-client
|
||||
```
|
||||
|
||||
LlamaIndex requires providing an instance of `QdrantClient`, so it can interact with Qdrant server.
|
||||
|
||||
```python
|
||||
from llama_index.vector_stores.qdrant import QdrantVectorStore
|
||||
|
||||
import qdrant_client
|
||||
|
||||
client = qdrant_client.QdrantClient(
|
||||
"<qdrant-url>",
|
||||
api_key="<qdrant-api-key>", # For Qdrant Cloud, None for local instance
|
||||
)
|
||||
|
||||
index = QdrantVectorStore(client=client, collection_name="documents")
|
||||
```
|
||||
|
||||
The library [comes with a notebook](https://github.com/jerryjliu/llama_index/blob/main/docs/examples/vector_stores/QdrantIndexDemo.ipynb)
|
||||
that shows an end-to-end example of how to use Qdrant within LlamaIndex.
|
||||
@@ -0,0 +1,96 @@
|
||||
---
|
||||
title: MindsDB
|
||||
weight: 1100
|
||||
---
|
||||
|
||||
# MindsDB
|
||||
|
||||
[MindsDB](https://mindsdb.com) is an AI automation platform for building AI/ML powered features and applications. It works by connecting any source of data with any AI/ML model or framework and automating how real-time data flows between them.
|
||||
|
||||
With the MindsDB-Qdrant integration, you can now select Qdrant as a database to load into and retrieve from with semantic search and filtering.
|
||||
|
||||
**MindsDB allows you to easily**:
|
||||
|
||||
- Connect to any store of data or end-user application.
|
||||
- Pass data to an AI model from any store of data or end-user application.
|
||||
- Plug the output of an AI model into any store of data or end-user application.
|
||||
- Fully automate these workflows to build AI-powered features and applications
|
||||
|
||||
## Usage
|
||||
|
||||
To get started with Qdrant and MindsDB, the following syntax can be used.
|
||||
|
||||
```sql
|
||||
CREATE DATABASE qdrant_test
|
||||
WITH ENGINE = "qdrant",
|
||||
PARAMETERS = {
|
||||
"location": ":memory:",
|
||||
"collection_config": {
|
||||
"size": 386,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The available arguments for instantiating Qdrant can be found [here](https://github.com/mindsdb/mindsdb/blob/23a509cb26bacae9cc22475497b8644e3f3e23c3/mindsdb/integrations/handlers/qdrant_handler/qdrant_handler.py#L408-L468).
|
||||
|
||||
## Creating a new table
|
||||
|
||||
- Qdrant options for creating a collection can be specified as `collection_config` in the `CREATE DATABASE` parameters.
|
||||
- By default, UUIDs are set as collection IDs. You can provide your own IDs under the `id` column.
|
||||
|
||||
```sql
|
||||
CREATE TABLE qdrant_test.test_table (
|
||||
SELECT embeddings,'{"source": "bbc"}' as metadata FROM mysql_demo_db.test_embeddings
|
||||
);
|
||||
```
|
||||
|
||||
## Querying the database
|
||||
|
||||
#### Perform a full retrieval using the following syntax.
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
```
|
||||
|
||||
By default, the `LIMIT` is set to 10 and the `OFFSET` is set to 0.
|
||||
|
||||
#### Perform a similarity search using your embeddings
|
||||
|
||||
<aside role="status">Qdrant supports <a href="https://qdrant.tech/documentation/concepts/indexing/#payload-index">payload indexing</a> that vastly improves retrieval efficiency with filters and is highly recommended. Please note that this feature currently cannot be configured via MindsDB and must be set up separately if needed.</aside>
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
WHERE search_vector = (select embeddings from mysql_demo_db.test_embeddings limit 1)
|
||||
```
|
||||
|
||||
#### Perform a search using filters
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
WHERE `metadata.source` = 'bbc';
|
||||
```
|
||||
|
||||
#### Delete entries using IDs
|
||||
|
||||
```sql
|
||||
DELETE FROM qtest.test_table_6
|
||||
WHERE id = 2
|
||||
```
|
||||
|
||||
#### Delete entries using filters
|
||||
|
||||
```sql
|
||||
DELETE * FROM qdrant_test.test_table
|
||||
WHERE `metadata.source` = 'bbc';
|
||||
```
|
||||
|
||||
#### Drop a table
|
||||
|
||||
```sql
|
||||
DROP TABLE qdrant_test.test_table;
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
You can find more information pertaining to MindsDB and its datasources [here](https://docs.mindsdb.com/).
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
title: PrivateGPT
|
||||
weight: 1600
|
||||
---
|
||||
|
||||
# PrivateGPT
|
||||
|
||||
[PrivateGPT](https://docs.privategpt.dev/) is a production-ready AI project that allows you to inquire about your documents using Large Language Models (LLMs) with offline support.
|
||||
|
||||
PrivateGPT uses Qdrant as the default vectorstore for ingesting and retrieving documents.
|
||||
|
||||
## Configuration
|
||||
|
||||
Qdrant settings can be configured by setting values to the qdrant property in the `settings.yaml` file. By default, Qdrant tries to connect to an instance at http://localhost:3000.
|
||||
|
||||
Example:
|
||||
```yaml
|
||||
qdrant:
|
||||
url: "https://xyz-example.eu-central.aws.cloud.qdrant.io:6333"
|
||||
api_key: "<your-api-key>"
|
||||
```
|
||||
|
||||
The available [configuration options](https://docs.privategpt.dev/manual/storage/vector-stores#qdrant-configuration) are:
|
||||
| Field | Description |
|
||||
|--------------|-------------|
|
||||
| location | If `:memory:` - use in-memory Qdrant instance.<br>If `str` - use it as a `url` parameter.|
|
||||
| url | Either host or str of `Optional[scheme], host, Optional[port], Optional[prefix]`.<br> Eg. `http://localhost:6333` |
|
||||
| port | Port of the REST API interface. Default: `6333` |
|
||||
| grpc_port | Port of the gRPC interface. Default: `6334` |
|
||||
| prefer_grpc | If `true` - use gRPC interface whenever possible in custom methods. |
|
||||
| https | If `true` - use HTTPS(SSL) protocol.|
|
||||
| api_key | API key for authentication in Qdrant Cloud.|
|
||||
| prefix | If set, add `prefix` to the REST URL path.<br>Example: `service/v1` will result in `http://localhost:6333/service/v1/{qdrant-endpoint}` for REST API.|
|
||||
| timeout | Timeout for REST and gRPC API requests.<br>Default: 5.0 seconds for REST and unlimited for gRPC |
|
||||
| host | Host name of Qdrant service. If url and host are not set, defaults to 'localhost'.|
|
||||
| path | Persistence path for QdrantLocal. Eg. `local_data/private_gpt/qdrant`|
|
||||
| force_disable_check_same_thread | Force disable check_same_thread for QdrantLocal sqlite connection.|
|
||||
|
||||
## Next steps
|
||||
|
||||
Find the PrivateGPT docs [here](https://docs.privategpt.dev/).
|
||||
@@ -0,0 +1,147 @@
|
||||
---
|
||||
title: Apache Spark
|
||||
weight: 1400
|
||||
---
|
||||
|
||||
# Apache Spark
|
||||
|
||||
[Spark](https://spark.apache.org/) is a leading distributed computing framework that empowers you to work with massive datasets efficiently. When it comes to leveraging the power of Spark for your data processing needs, the [Qdrant-Spark Connector](https://github.com/qdrant/qdrant-spark) is to be considered. This connector enables Qdrant to serve as a storage destination in Spark, offering a seamless bridge between the two.
|
||||
|
||||
## Installation
|
||||
|
||||
You can set up the Qdrant-Spark Connector in a few different ways, depending on your preferences and requirements.
|
||||
|
||||
### GitHub Releases
|
||||
|
||||
The simplest way to get started is by downloading pre-packaged JAR file releases from the [Qdrant-Spark GitHub releases page](https://github.com/qdrant/qdrant-spark/releases). These JAR files come with all the necessary dependencies to get you going.
|
||||
|
||||
### Building from Source
|
||||
|
||||
If you prefer to build the JAR from source, you'll need [JDK 17](https://www.oracle.com/java/technologies/javase/jdk17-archive-downloads.html) and [Maven](https://maven.apache.org/) installed on your system. Once you have the prerequisites in place, navigate to the project's root directory and run the following command:
|
||||
|
||||
```bash
|
||||
mvn package -P assembly
|
||||
```
|
||||
This command will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default.
|
||||
|
||||
### Maven Central
|
||||
|
||||
For Java and Scala projects, you can also obtain the Qdrant-Spark Connector from [Maven Central](https://central.sonatype.com/artifact/io.qdrant/spark).
|
||||
|
||||
```xml
|
||||
<dependency>
|
||||
<groupId>io.qdrant</groupId>
|
||||
<artifactId>spark</artifactId>
|
||||
<version>1.6</version>
|
||||
</dependency>
|
||||
```
|
||||
|
||||
## Getting Started
|
||||
|
||||
After successfully installing the Qdrant-Spark Connector, you can start integrating Qdrant with your Spark applications. Below, we'll walk through the basic steps of creating a Spark session with Qdrant support and loading data into Qdrant.
|
||||
|
||||
### Creating a single-node Spark session with Qdrant Support
|
||||
|
||||
To begin, import the necessary libraries and create a Spark session with Qdrant support. Here's how:
|
||||
|
||||
```python
|
||||
from pyspark.sql import SparkSession
|
||||
|
||||
spark = SparkSession.builder.config(
|
||||
"spark.jars",
|
||||
"spark-1.0-assembly.jar", # Specify the downloaded JAR file
|
||||
)
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate()
|
||||
```
|
||||
|
||||
```scala
|
||||
import org.apache.spark.sql.SparkSession
|
||||
|
||||
val spark = SparkSession.builder
|
||||
.config("spark.jars", "spark-1.0-assembly.jar") // Specify the downloaded JAR file
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate()
|
||||
```
|
||||
|
||||
```java
|
||||
import org.apache.spark.sql.SparkSession;
|
||||
|
||||
public class QdrantSparkJavaExample {
|
||||
public static void main(String[] args) {
|
||||
SparkSession spark = SparkSession.builder()
|
||||
.config("spark.jars", "spark-1.0-assembly.jar") // Specify the downloaded JAR file
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate();
|
||||
...
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Loading Data into Qdrant
|
||||
|
||||
<aside role="status">To load data into Qdrant, you'll need to create a collection with the appropriate vector dimensions and configurations in advance.</aside>
|
||||
|
||||
Here's how you can use the Qdrant-Spark Connector to upsert data:
|
||||
|
||||
```python
|
||||
<YourDataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", <QDRANT_URL>) # REST URL of the Qdrant instance
|
||||
.option("collection_name", <QDRANT_COLLECTION_NAME>) # Name of the collection to write data into
|
||||
.option("embedding_field", <EMBEDDING_FIELD_NAME>) # Name of the field holding the embeddings
|
||||
.option("schema", <YourDataFrame>.schema.json()) # JSON string of the dataframe schema
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
```scala
|
||||
<YourDataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", QDRANT_URL) // REST URL of the Qdrant instance
|
||||
.option("collection_name", QDRANT_COLLECTION_NAME) // Name of the collection to write data into
|
||||
.option("embedding_field", EMBEDDING_FIELD_NAME) // Name of the field holding the embeddings
|
||||
.option("schema", <YourDataFrame>.schema.json()) // JSON string of the dataframe schema
|
||||
.mode("append")
|
||||
.save()
|
||||
|
||||
```
|
||||
|
||||
```java
|
||||
<YourDataFrame>
|
||||
.write()
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", QDRANT_URL) // REST URL of the Qdrant instance
|
||||
.option("collection_name", QDRANT_COLLECTION_NAME) // Name of the collection to write data into
|
||||
.option("embedding_field", EMBEDDING_FIELD_NAME) // Name of the field holding the embeddings
|
||||
.option("schema", <YourDataFrame>.schema().json()) // JSON string of the dataframe schema
|
||||
.mode("append")
|
||||
.save();
|
||||
```
|
||||
|
||||
## Datatype Support
|
||||
|
||||
Qdrant supports all the Spark data types, and the appropriate data types are mapped based on the provided schema.
|
||||
|
||||
## Options and Spark Types
|
||||
|
||||
The Qdrant-Spark Connector provides a range of options to fine-tune your data integration process. Here's a quick reference:
|
||||
|
||||
| Option | Description | DataType | Required |
|
||||
| :---------------- | :--------------------------------------------------------------------------- | :--------------------- | :------- |
|
||||
| `qdrant_url` | REST URL of the Qdrant instance | `StringType` | ✅ |
|
||||
| `collection_name` | Name of the collection to write data into | `StringType` | ✅ |
|
||||
| `embedding_field` | Name of the field holding the embeddings | `ArrayType(FloatType)` | ✅ |
|
||||
| `schema` | JSON string of the dataframe schema | `StringType` | ✅ |
|
||||
| `mode` | Write mode of the dataframe | `StringType` | ✅ |
|
||||
| `id_field` | Name of the field holding the point IDs. Default: A random UUID is generated | `StringType` | ❌ |
|
||||
| `batch_size` | Max size of the upload batch. Default: 100 | `IntType` | ❌ |
|
||||
| `retries` | Number of upload retries. Default: 3 | `IntType` | ❌ |
|
||||
| `api_key` | Qdrant API key for authenticated requests. Default: null | `StringType` | ❌ |
|
||||
|
||||
For more information, be sure to check out the [Qdrant-Spark GitHub repository](https://github.com/qdrant/qdrant-spark). The Apache Spark guide is available [here](https://spark.apache.org/docs/latest/quick-start.html). Happy data processing!
|
||||
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: txtai
|
||||
weight: 500
|
||||
---
|
||||
|
||||
# txtai
|
||||
|
||||
Qdrant might be also used as an embedding backend in [txtai](https://neuml.github.io/txtai/) semantic applications.
|
||||
|
||||
txtai simplifies building AI-powered semantic search applications using Transformers. It leverages the neural embeddings and their
|
||||
properties to encode high-dimensional data in a lower-dimensional space and allows to find similar objects based on their embeddings'
|
||||
proximity.
|
||||
|
||||
Qdrant is not built-in txtai backend and requires installing an additional dependency:
|
||||
|
||||
```bash
|
||||
pip install qdrant-txtai
|
||||
```
|
||||
|
||||
The examples and some more information might be found in [qdrant-txtai repository](https://github.com/qdrant/qdrant-txtai).
|
||||
|
||||
|
||||
Reference in New Issue
Block a user