mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-29 07:58:31 +02:00
docs: Restructured integrations section (#1089)
* docs: Reorder integrations * docs: Formatting langchain-go.md * docs: Title for index * docs: Redpanda docs (#1092)
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
---
|
||||
title: Data Management
|
||||
weight: 15
|
||||
---
|
||||
|
||||
## Data Management Integrations
|
||||
|
||||
| Integration | Description |
|
||||
| ------------------------------- | -------------------------------------------------------------------------------------------------- |
|
||||
| [Airbyte](./airbyte/) | Data integration platform specialising in ELT pipelines. |
|
||||
| [Airflow](./airflow/) | Platform designed for developing, scheduling, and monitoring batch-oriented workflows. |
|
||||
| [Redpanda Connect](./redpanda/) | Declarative data-agnostic streaming service for efficient, stateless processing. |
|
||||
| [Confluent](./confluent/) | Fully-managed data streaming platform with a cloud-native Apache Kafka engine. |
|
||||
| [DLT](./dlt/) | Python library to simplify data loading processes between several sources and destinations. |
|
||||
| [Fondant](./fondant/) | Framework for developing datasets, sharing reusable operations and data processing trees. |
|
||||
| [MindsDB](./mindsdb/) | Platform to deploy, serve, and fine-tune models with numerous data source integrations. |
|
||||
| [Apache NiFi](./nifi/) | Data ingestion platform to manage data transfer between different sources and destination systems. |
|
||||
| [Apache Spark](./spark/) | A unified analytics engine for large-scale data processing. |
|
||||
| [Unstructured](./unstructured/) | Python library with components for ingesting and pre-processing data from numerous sources. |
|
||||
@@ -0,0 +1,79 @@
|
||||
---
|
||||
title: Airbyte
|
||||
aliases: [ ../integrations/airbyte/, ../frameworks/airbyte/ ]
|
||||
---
|
||||
|
||||
# Airbyte
|
||||
|
||||
[Airbyte](https://airbyte.com/) is an open-source data integration platform that helps you replicate your data
|
||||
between different systems. It has a [growing list of connectors](https://docs.airbyte.io/integrations) that can
|
||||
be used to ingest data from multiple sources. Building data pipelines is also crucial for managing the data in
|
||||
Qdrant, and Airbyte is a great tool for this purpose.
|
||||
|
||||
Airbyte may take care of the data ingestion from a selected source, while Qdrant will help you to build a search
|
||||
engine on top of it. There are three supported modes of how the data can be ingested into Qdrant:
|
||||
|
||||
* **Full Refresh Sync**
|
||||
* **Incremental - Append Sync**
|
||||
* **Incremental - Append + Deduped**
|
||||
|
||||
You can read more about these modes in the [Airbyte documentation](https://docs.airbyte.io/integrations/destinations/qdrant).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before you start, make sure you have the following:
|
||||
|
||||
1. Airbyte instance, either [Open Source](https://airbyte.com/solutions/airbyte-open-source),
|
||||
[Self-Managed](https://airbyte.com/solutions/airbyte-enterprise), or [Cloud](https://airbyte.com/solutions/airbyte-cloud).
|
||||
2. Running instance of Qdrant. It has to be accessible by URL from the machine where Airbyte is running.
|
||||
You can follow the [installation guide](/documentation/guides/installation/) to set up Qdrant.
|
||||
|
||||
## Setting up Qdrant as a destination
|
||||
|
||||
Once you have a running instance of Airbyte, you can set up Qdrant as a destination directly in the UI.
|
||||
Airbyte's Qdrant destination is connected with a single collection in Qdrant.
|
||||
|
||||

|
||||
|
||||
### Text processing
|
||||
|
||||
Airbyte has some built-in mechanisms to transform your texts into embeddings. You can choose how you want to
|
||||
chunk your fields into pieces before calculating the embeddings, but also which fields should be used to
|
||||
create the point payload.
|
||||
|
||||

|
||||
|
||||
### Embeddings
|
||||
|
||||
You can choose the model that will be used to calculate the embeddings. Currently, Airbyte supports multiple
|
||||
models, including OpenAI and Cohere.
|
||||
|
||||

|
||||
|
||||
Using some precomputed embeddings from your data source is also possible. In this case, you can pass the field
|
||||
name containing the embeddings and their dimensionality.
|
||||
|
||||

|
||||
|
||||
### Qdrant connection details
|
||||
|
||||
Finally, we can configure the target Qdrant instance and collection. In case you use the built-in authentication
|
||||
mechanism, here is where you can pass the token.
|
||||
|
||||

|
||||
|
||||
Once you confirm creating the destination, Airbyte will test if a specified Qdrant cluster is accessible and
|
||||
might be used as a destination.
|
||||
|
||||
## Setting up connection
|
||||
|
||||
Airbyte combines sources and destinations into a single entity called a connection. Once you have a destination
|
||||
configured and a source, you can create a connection between them. It doesn't matter what source you use, as
|
||||
long as Airbyte supports it. The process is pretty straightforward, but depends on the source you use.
|
||||
|
||||

|
||||
|
||||
## Further Reading
|
||||
|
||||
* [Airbyte documentation](https://docs.airbyte.com/understanding-airbyte/connections/).
|
||||
* [Source Code](https://github.com/airbytehq/airbyte/tree/master/airbyte-integrations/connectors/destination-qdrant)
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
title: Apache Airflow
|
||||
aliases: [ ../frameworks/airflow/ ]
|
||||
---
|
||||
|
||||
# Apache Airflow
|
||||
|
||||
[Apache Airflow](https://airflow.apache.org/) is an open-source platform for authoring, scheduling and monitoring data and computing workflows. Airflow uses Python to create workflows that can be easily scheduled and monitored.
|
||||
|
||||
Qdrant is available as a [provider](https://airflow.apache.org/docs/apache-airflow-providers-qdrant/stable/index.html) in Airflow to interface with the database.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before configuring Airflow, you need:
|
||||
|
||||
1. A Qdrant instance to connect to. You can set one up in our [installation guide](/documentation/guides/installation/).
|
||||
|
||||
2. A running Airflow instance. You can use their [Quick Start Guide](https://airflow.apache.org/docs/apache-airflow/stable/start.html).
|
||||
|
||||
## Installation
|
||||
|
||||
You can install the Qdrant provider by running `pip install apache-airflow-providers-qdrant` in your Airflow shell.
|
||||
|
||||
**NOTE**: You'll have to restart your Airflow session for the provider to be available.
|
||||
|
||||
## Setting up a connection
|
||||
|
||||
Open the `Admin-> Connections` section of the Airflow UI. Click the `Create` link to create a new [Qdrant connection](https://airflow.apache.org/docs/apache-airflow-providers-qdrant/stable/connections.html).
|
||||
|
||||

|
||||
|
||||
You can also set up a connection using [environment variables](https://airflow.apache.org/docs/apache-airflow/stable/howto/connection.html#environment-variables-connections) or an [external secret backend](https://airflow.apache.org/docs/apache-airflow/stable/security/secrets/secrets-backend/index.html).
|
||||
|
||||
## Qdrant hook
|
||||
|
||||
An Airflow hook is an abstraction of a specific API that allows Airflow to interact with an external system.
|
||||
|
||||
```python
|
||||
from airflow.providers.qdrant.hooks.qdrant import QdrantHook
|
||||
|
||||
hook = QdrantHook(conn_id="qdrant_connection")
|
||||
|
||||
hook.verify_connection()
|
||||
```
|
||||
|
||||
A [`qdrant_client#QdrantClient`](https://pypi.org/project/qdrant-client/) instance is available via `@property conn` of the `QdrantHook` instance for use within your Airflow workflows.
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
hook.conn.count("<COLLECTION_NAME>")
|
||||
|
||||
hook.conn.upsert(
|
||||
"<COLLECTION_NAME>",
|
||||
points=[
|
||||
models.PointStruct(id=32, vector=[0.32, 0.12, 0.123], payload={"color": "red"})
|
||||
],
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
## Qdrant Ingest Operator
|
||||
|
||||
The Qdrant provider also provides a convenience operator for uploading data to a Qdrant collection that internally uses the Qdrant hook.
|
||||
|
||||
```python
|
||||
from airflow.providers.qdrant.operators.qdrant import QdrantIngestOperator
|
||||
|
||||
vectors = [
|
||||
[0.11, 0.22, 0.33, 0.44],
|
||||
[0.55, 0.66, 0.77, 0.88],
|
||||
[0.88, 0.11, 0.12, 0.13],
|
||||
]
|
||||
ids = [32, 21, "b626f6a9-b14d-4af9-b7c3-43d8deb719a6"]
|
||||
payload = [{"meta": "data"}, {"meta": "data_2"}, {"meta": "data_3", "extra": "data"}]
|
||||
|
||||
QdrantIngestOperator(
|
||||
conn_id="qdrant_connection",
|
||||
task_id="qdrant_ingest",
|
||||
collection_name="<COLLECTION_NAME>",
|
||||
vectors=vectors,
|
||||
ids=ids,
|
||||
payload=payload,
|
||||
)
|
||||
```
|
||||
|
||||
## Reference
|
||||
- 📦 [Provider package PyPI](https://pypi.org/project/apache-airflow-providers-qdrant/)
|
||||
- 📚 [Provider docs](https://airflow.apache.org/docs/apache-airflow-providers-qdrant/stable/index.html)
|
||||
- 📄 [Source Code](https://github.com/apache/airflow/tree/main/airflow/providers/qdrant)
|
||||
@@ -0,0 +1,283 @@
|
||||
---
|
||||
title: Confluent Kafka
|
||||
aliases: [ ../frameworks/confluent/ ]
|
||||
---
|
||||
|
||||

|
||||
|
||||
Built by the original creators of Apache Kafka®, [Confluent Cloud](https://www.confluent.io/confluent-cloud/?utm_campaign=tm.pmm_cd.cwc_partner_Qdrant_generic&utm_source=Qdrant&utm_medium=partnerref) is a cloud-native and complete data streaming platform available on AWS, Azure, and Google Cloud. The platform includes a fully managed, elastically scaling Kafka engine, 120+ connectors, serverless Apache Flink®, enterprise-grade security controls, and a robust governance suite.
|
||||
|
||||
With our [Qdrant-Kafka Sink Connector](https://github.com/qdrant/qdrant-kafka), Qdrant is part of the [Connect with Confluent](https://www.confluent.io/partners/connect/) technology partner program. It brings fully managed data streams directly to organizations from Confluent Cloud, making it easier for organizations to stream any data to Qdrant with a fully managed Apache Kafka service.
|
||||
|
||||
## Usage
|
||||
|
||||
### Pre-requisites
|
||||
|
||||
- A Confluent Cloud account. You can begin with a [free trial](https://www.confluent.io/confluent-cloud/tryfree/?utm_campaign=tm.pmm_cd.cwc_partner_qdrant_tryfree&utm_source=qdrant&utm_medium=partnerref) with credits for the first 30 days.
|
||||
- Qdrant instance to connect to. You can get a free cloud instance at [cloud.qdrant.io](https://cloud.qdrant.io/).
|
||||
|
||||
### Installation
|
||||
|
||||
1) Download the latest connector zip file from [Confluent Hub](https://www.confluent.io/hub/qdrant/qdrant-kafka).
|
||||
|
||||
2) Configure an environment and cluster on Confluent and create a topic to produce messages for.
|
||||
|
||||
3) Navigate to the `Connectors` section of the Confluent cluster and click `Add Plugin`. Upload the zip file with the following info.
|
||||
|
||||

|
||||
|
||||
4) Once installed, navigate to the connector and set the following configuration values.
|
||||
|
||||

|
||||
|
||||
Replace the placeholder values with your credentials.
|
||||
|
||||
5) Add the Qdrant instance host to the allowed networking endpoints.
|
||||
|
||||

|
||||
|
||||
7) Start the connector.
|
||||
|
||||
## Producing Messages
|
||||
|
||||
You can now produce messages for the configured topic, and they'll be written into the configured Qdrant instance.
|
||||
|
||||

|
||||
|
||||
## Message Formats
|
||||
|
||||
The connector supports messages in the following formats.
|
||||
|
||||
_Click each to expand._
|
||||
|
||||
<details>
|
||||
<summary><b>Unnamed/Default vector</b></summary>
|
||||
|
||||
Reference: [Creating a collection with a default vector](https://qdrant.tech/documentation/concepts/collections/#create-a-collection).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
],
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Named multiple vectors</b></summary>
|
||||
|
||||
Reference: [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"some-dense": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
],
|
||||
"some-other-dense": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
]
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Sparse vectors</b></summary>
|
||||
|
||||
Reference: [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"some-sparse": {
|
||||
"indices": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9
|
||||
],
|
||||
"values": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8,
|
||||
0.9,
|
||||
1.0
|
||||
]
|
||||
}
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Multi-vectors</b></summary>
|
||||
|
||||
Reference:
|
||||
|
||||
- [Multi-vectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors)
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"some-multi": [
|
||||
[
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8,
|
||||
0.9,
|
||||
1.0
|
||||
],
|
||||
[
|
||||
1.0,
|
||||
0.9,
|
||||
0.8,
|
||||
0.5,
|
||||
0.4,
|
||||
0.8,
|
||||
0.6,
|
||||
0.4,
|
||||
0.2,
|
||||
0.1
|
||||
]
|
||||
]
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Combination of named dense and sparse vectors</b></summary>
|
||||
|
||||
Reference:
|
||||
|
||||
- [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors).
|
||||
|
||||
- [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": "a10435b5-2a58-427a-a3a0-a5d845b147b7",
|
||||
"vector": {
|
||||
"some-other-dense": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
],
|
||||
"some-sparse": {
|
||||
"indices": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9
|
||||
],
|
||||
"values": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8,
|
||||
0.9,
|
||||
1.0
|
||||
]
|
||||
}
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Kafka Connect Docs](https://docs.confluent.io/platform/current/connect/index.html)
|
||||
- [Confluent Connectors Docs](https://docs.confluent.io/cloud/current/connectors/bring-your-connector/custom-connector-qs.html)
|
||||
@@ -0,0 +1,102 @@
|
||||
---
|
||||
title: DLT
|
||||
aliases: [ ../integrations/dlt/, ../frameworks/dlt/ ]
|
||||
---
|
||||
|
||||
# DLT(Data Load Tool)
|
||||
|
||||
[DLT](https://dlthub.com/) is an open-source library that you can add to your Python scripts to load data from various and often messy data sources into well-structured, live datasets.
|
||||
|
||||
With the DLT-Qdrant integration, you can now select Qdrant as a DLT destination to load data into.
|
||||
|
||||
**DLT Enables**
|
||||
|
||||
- Automated maintenance - with schema inference, alerts and short declarative code, maintenance becomes simple.
|
||||
- Run it where Python runs - on Airflow, serverless functions, notebooks. Scales on micro and large infrastructure alike.
|
||||
- User-friendly, declarative interface that removes knowledge obstacles for beginners while empowering senior professionals.
|
||||
|
||||
## Usage
|
||||
|
||||
To get started, install `dlt` with the `qdrant` extra.
|
||||
|
||||
```bash
|
||||
pip install "dlt[qdrant]"
|
||||
```
|
||||
|
||||
Configure the destination in the DLT secrets file. The file is located at `~/.dlt/secrets.toml` by default. Add the following section to the secrets file.
|
||||
|
||||
```toml
|
||||
[destination.qdrant.credentials]
|
||||
location = "https://your-qdrant-url"
|
||||
api_key = "your-qdrant-api-key"
|
||||
```
|
||||
|
||||
The location will default to `http://localhost:6333` and `api_key` is not defined - which are the defaults for a local Qdrant instance.
|
||||
Find more information about DLT configurations [here](https://dlthub.com/docs/general-usage/credentials).
|
||||
|
||||
Define the source of the data.
|
||||
|
||||
```python
|
||||
import dlt
|
||||
from dlt.destinations.qdrant import qdrant_adapter
|
||||
|
||||
movies = [
|
||||
{
|
||||
"title": "Blade Runner",
|
||||
"year": 1982,
|
||||
"description": "The film is about a dystopian vision of the future that combines noir elements with sci-fi imagery."
|
||||
},
|
||||
{
|
||||
"title": "Ghost in the Shell",
|
||||
"year": 1995,
|
||||
"description": "The film is about a cyborg policewoman and her partner who set out to find the main culprit behind brain hacking, the Puppet Master."
|
||||
},
|
||||
{
|
||||
"title": "The Matrix",
|
||||
"year": 1999,
|
||||
"description": "The movie is set in the 22nd century and tells the story of a computer hacker who joins an underground group fighting the powerful computers that rule the earth."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
<aside role="status">
|
||||
A more comprehensive pipeline would load data from some API or use one of <a href="https://dlthub.com/docs/dlt-ecosystem/verified-sources">DLT's verified sources</a>.
|
||||
</aside>
|
||||
|
||||
Define the pipeline.
|
||||
|
||||
```python
|
||||
pipeline = dlt.pipeline(
|
||||
pipeline_name="movies",
|
||||
destination="qdrant",
|
||||
dataset_name="movies_dataset",
|
||||
)
|
||||
```
|
||||
|
||||
Run the pipeline.
|
||||
|
||||
```python
|
||||
info = pipeline.run(
|
||||
qdrant_adapter(
|
||||
movies,
|
||||
embed=["title", "description"]
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
The data is now loaded into Qdrant.
|
||||
|
||||
To use vector search after the data has been loaded, you must specify which fields Qdrant needs to generate embeddings for. You do that by wrapping the data (or [DLT resource](https://dlthub.com/docs/general-usage/resource)) with the `qdrant_adapter` function.
|
||||
|
||||
## Write disposition
|
||||
|
||||
A DLT [write disposition](https://dlthub.com/docs/dlt-ecosystem/destinations/qdrant/#write-disposition) defines how the data should be written to the destination. All write dispositions are supported by the Qdrant destination.
|
||||
|
||||
## DLT Sync
|
||||
|
||||
Qdrant destination supports syncing of the [`DLT` state](https://dlthub.com/docs/general-usage/state#syncing-state-with-destination).
|
||||
|
||||
## Next steps
|
||||
|
||||
- The comprehensive Qdrant DLT destination documentation can be found [here](https://dlthub.com/docs/dlt-ecosystem/destinations/qdrant/).
|
||||
- [Source Code](https://github.com/dlt-hub/dlt/tree/devel/dlt/destinations/impl/qdrant)
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
title: Fondant
|
||||
aliases: [ ../integrations/fondant/, ../frameworks/fondant/ ]
|
||||
---
|
||||
|
||||
# Fondant
|
||||
|
||||
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed
|
||||
up large-scale data processing by making containerized components reusable across pipelines and
|
||||
execution environments. Benefit from built-in features such as autoscaling, data lineage, and
|
||||
pipeline caching, and deploy to (managed) platforms such as Vertex AI, Sagemaker, and Kubeflow
|
||||
Pipelines.
|
||||
|
||||
Fondant comes with a library of reusable components that you can leverage to compose your own
|
||||
pipeline, including a Qdrant component for writing embeddings to Qdrant.
|
||||
|
||||
## Usage
|
||||
|
||||
<aside role="status">
|
||||
A Qdrant collection has to be <a href="/documentation/concepts/collections/">created in advance</a>
|
||||
</aside>
|
||||
|
||||
**A data load pipeline for RAG using Qdrant**.
|
||||
|
||||
A simple ingestion pipeline could look like the following:
|
||||
|
||||
```python
|
||||
import pyarrow as pa
|
||||
from fondant.pipeline import Pipeline
|
||||
|
||||
indexing_pipeline = Pipeline(
|
||||
name="ingestion-pipeline",
|
||||
description="Pipeline to prepare and process data for building a RAG solution",
|
||||
base_path="./fondant-artifacts",
|
||||
)
|
||||
|
||||
# An custom implemenation of a read component.
|
||||
text = indexing_pipeline.read(
|
||||
"path/to/data-source-component",
|
||||
arguments={
|
||||
# your custom arguments
|
||||
}
|
||||
)
|
||||
|
||||
chunks = text.apply(
|
||||
"chunk_text",
|
||||
arguments={
|
||||
"chunk_size": 512,
|
||||
"chunk_overlap": 32,
|
||||
},
|
||||
)
|
||||
|
||||
embeddings = chunks.apply(
|
||||
"embed_text",
|
||||
arguments={
|
||||
"model_provider": "huggingface",
|
||||
"model": "all-MiniLM-L6-v2",
|
||||
},
|
||||
)
|
||||
|
||||
embeddings.write(
|
||||
"index_qdrant",
|
||||
arguments={
|
||||
"url": "http:localhost:6333",
|
||||
"collection_name": "some-collection-name",
|
||||
},
|
||||
cache=False,
|
||||
)
|
||||
```
|
||||
|
||||
Once you have a pipeline, you can easily run it using the built-in CLI. Fondant allows
|
||||
you to run the pipeline in production across different clouds.
|
||||
|
||||
The first component is a custom read module that needs to be implemented and cannot be used off the
|
||||
shelf. A detailed tutorial on how to rebuild this
|
||||
pipeline [is provided on GitHub](https://github.com/ml6team/fondant-usecase-RAG/tree/main).
|
||||
|
||||
## Next steps
|
||||
|
||||
More information about creating your own pipelines and components can be found in the [Fondant
|
||||
documentation](https://fondant.ai/en/stable/).
|
||||
@@ -0,0 +1,97 @@
|
||||
---
|
||||
title: MindsDB
|
||||
aliases: [ ../integrations/mindsdb/, ../frameworks/mindsdb/ ]
|
||||
---
|
||||
|
||||
# MindsDB
|
||||
|
||||
[MindsDB](https://mindsdb.com) is an AI automation platform for building AI/ML powered features and applications. It works by connecting any source of data with any AI/ML model or framework and automating how real-time data flows between them.
|
||||
|
||||
With the MindsDB-Qdrant integration, you can now select Qdrant as a database to load into and retrieve from with semantic search and filtering.
|
||||
|
||||
**MindsDB allows you to easily**:
|
||||
|
||||
- Connect to any store of data or end-user application.
|
||||
- Pass data to an AI model from any store of data or end-user application.
|
||||
- Plug the output of an AI model into any store of data or end-user application.
|
||||
- Fully automate these workflows to build AI-powered features and applications
|
||||
|
||||
## Usage
|
||||
|
||||
To get started with Qdrant and MindsDB, the following syntax can be used.
|
||||
|
||||
```sql
|
||||
CREATE DATABASE qdrant_test
|
||||
WITH ENGINE = "qdrant",
|
||||
PARAMETERS = {
|
||||
"location": ":memory:",
|
||||
"collection_config": {
|
||||
"size": 386,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The available arguments for instantiating Qdrant can be found [here](https://github.com/mindsdb/mindsdb/blob/23a509cb26bacae9cc22475497b8644e3f3e23c3/mindsdb/integrations/handlers/qdrant_handler/qdrant_handler.py#L408-L468).
|
||||
|
||||
## Creating a new table
|
||||
|
||||
- Qdrant options for creating a collection can be specified as `collection_config` in the `CREATE DATABASE` parameters.
|
||||
- By default, UUIDs are set as collection IDs. You can provide your own IDs under the `id` column.
|
||||
|
||||
```sql
|
||||
CREATE TABLE qdrant_test.test_table (
|
||||
SELECT embeddings,'{"source": "bbc"}' as metadata FROM mysql_demo_db.test_embeddings
|
||||
);
|
||||
```
|
||||
|
||||
## Querying the database
|
||||
|
||||
#### Perform a full retrieval using the following syntax.
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
```
|
||||
|
||||
By default, the `LIMIT` is set to 10 and the `OFFSET` is set to 0.
|
||||
|
||||
#### Perform a similarity search using your embeddings
|
||||
|
||||
<aside role="status">Qdrant supports <a href="/documentation/concepts/indexing/#payload-index">payload indexing</a> that vastly improves retrieval efficiency with filters and is highly recommended. Please note that this feature currently cannot be configured via MindsDB and must be set up separately if needed.</aside>
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
WHERE search_vector = (select embeddings from mysql_demo_db.test_embeddings limit 1)
|
||||
```
|
||||
|
||||
#### Perform a search using filters
|
||||
|
||||
```sql
|
||||
SELECT * FROM qdrant_test.test_table
|
||||
WHERE `metadata.source` = 'bbc';
|
||||
```
|
||||
|
||||
#### Delete entries using IDs
|
||||
|
||||
```sql
|
||||
DELETE FROM qtest.test_table_6
|
||||
WHERE id = 2
|
||||
```
|
||||
|
||||
#### Delete entries using filters
|
||||
|
||||
```sql
|
||||
DELETE * FROM qdrant_test.test_table
|
||||
WHERE `metadata.source` = 'bbc';
|
||||
```
|
||||
|
||||
#### Drop a table
|
||||
|
||||
```sql
|
||||
DROP TABLE qdrant_test.test_table;
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- You can find more information pertaining to MindsDB and its datasources [here](https://docs.mindsdb.com/).
|
||||
- [Source Code](https://github.com/mindsdb/mindsdb/tree/main/mindsdb/integrations/handlers/qdrant_handler)
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
title: Apache NiFi
|
||||
aliases: [ ../frameworks/nifi/ ]
|
||||
---
|
||||
|
||||
# Apache NiFi
|
||||
|
||||
[NiFi](https://nifi.apache.org/) is a real-time data ingestion platform, which can transfer and manage data transfer between numerous sources and destination systems. It supports many protocols and offers a web-based user interface for developing and monitoring data flows.
|
||||
|
||||
NiFi supports ingesting and querying data in Qdrant via its processor modules.
|
||||
|
||||
## Configuration
|
||||
|
||||

|
||||
|
||||
You can configure Qdrant NiFi processors with your Qdrant credentials, query/upload configurations. The processors offer 2 built-in embedding providers to encode data into vector embeddings - HuggingFace, OpenAI.
|
||||
|
||||
## Put Qdrant
|
||||
|
||||

|
||||
|
||||
The `Put Qdrant` processor can ingest NiFi [FlowFile](https://nifi.apache.org/docs/nifi-docs/html/nifi-in-depth.html#intro) data into a Qdrant collection.
|
||||
|
||||
## Query Qdrant
|
||||
|
||||

|
||||
|
||||
The `Query Qdrant` processor can perform a similarity search across a Qdrant collection and return a [FlowFile](https://nifi.apache.org/docs/nifi-docs/html/nifi-in-depth.html#intro) result.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [NiFi Documentation](https://nifi.apache.org/documentation/v2/).
|
||||
- [Source Code](https://github.com/apache/nifi-python-extensions)
|
||||
@@ -0,0 +1,47 @@
|
||||
---
|
||||
title: Redpanda Connect
|
||||
---
|
||||
|
||||
[Redpanda Connect](https://www.redpanda.com/connect) is a declarative data-agnostic streaming service designed for efficient, stateless processing steps. It offers transaction-based resiliency with back pressure, ensuring at-least-once delivery when connecting to at-least-once sources with sinks, without the need to persist messages during transit.
|
||||
|
||||
Connect pipelines are configured using a YAML file, which organizes components hierarchically. Each section represents a different component type, such as inputs, processors and outputs, and these can have nested child components and [dynamic values](https://docs.redpanda.com/redpanda-connect/configuration/interpolation/).
|
||||
|
||||
The [Qdrant Output](https://docs.redpanda.com/redpanda-connect/components/outputs/qdrant/) component enables streaming vector data into Qdrant collections in your RedPanda pipelines.
|
||||
|
||||
## Example
|
||||
|
||||
An example configuration of the output once the inputs and processors are set, would look like:
|
||||
|
||||
```yaml
|
||||
input:
|
||||
# https://docs.redpanda.com/redpanda-connect/components/inputs/about/
|
||||
|
||||
pipeline:
|
||||
processors:
|
||||
# https://docs.redpanda.com/redpanda-connect/components/processors/about/
|
||||
|
||||
output:
|
||||
label: "qdrant-output"
|
||||
qdrant:
|
||||
max_in_flight: 64
|
||||
batching:
|
||||
count: 8
|
||||
grpc_host: xyz-example.eu-central.aws.cloud.qdrant.io:6334
|
||||
api_token: "<provide-your-own-key>"
|
||||
tls:
|
||||
enabled: true
|
||||
# skip_cert_verify: false
|
||||
# enable_renegotiation: false
|
||||
# root_cas: ""
|
||||
# root_cas_file: ""
|
||||
# client_certs: []
|
||||
collection_name: "<collection_name>"
|
||||
id: root = uuid_v4()
|
||||
vector_mapping: 'root = {"some_dense": this.vector, "some_sparse": {"indices": [23,325,532],"values": [0.352,0.532,0.532]}}'
|
||||
payload_mapping: 'root = {"field": this.value, "field_2": 987}'
|
||||
```
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Getting started with Connect](https://docs.redpanda.com/redpanda-connect/guides/getting_started/)
|
||||
- [Qdrant Output Reference](https://docs.redpanda.com/redpanda-connect/components/outputs/qdrant/)
|
||||
@@ -0,0 +1,254 @@
|
||||
---
|
||||
title: Apache Spark
|
||||
aliases: [ ../integrations/spark/, ../frameworks/spark/ ]
|
||||
---
|
||||
|
||||
# Apache Spark
|
||||
|
||||
[Spark](https://spark.apache.org/) is a distributed computing framework designed for big data processing and analytics. The [Qdrant-Spark connector](https://github.com/qdrant/qdrant-spark) enables Qdrant to be a storage destination in Spark.
|
||||
|
||||
## Installation
|
||||
|
||||
You can set up the Qdrant-Spark Connector in a few different ways, depending on your preferences and requirements.
|
||||
|
||||
### GitHub Releases
|
||||
|
||||
The simplest way to get started is by downloading pre-packaged JAR file releases from the [GitHub releases page](https://github.com/qdrant/qdrant-spark/releases). These JAR files come with all the necessary dependencies.
|
||||
|
||||
### Building from Source
|
||||
|
||||
If you prefer to build the JAR from source, you'll need [JDK 8](https://www.azul.com/downloads/#zulu) and [Maven](https://maven.apache.org/) installed on your system. Once you have the prerequisites in place, navigate to the project's root directory and run the following command:
|
||||
|
||||
```bash
|
||||
mvn package
|
||||
```
|
||||
|
||||
This command will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default.
|
||||
|
||||
### Maven Central
|
||||
|
||||
For use with Java and Scala projects, the package can be found [here](https://central.sonatype.com/artifact/io.qdrant/spark).
|
||||
|
||||
## Usage
|
||||
|
||||
Below, we'll walk through the steps of creating a Spark session with Qdrant support and loading data into Qdrant.
|
||||
|
||||
### Creating a single-node Spark session with Qdrant Support
|
||||
|
||||
To begin, import the necessary libraries and create a Spark session with Qdrant support:
|
||||
|
||||
```python
|
||||
from pyspark.sql import SparkSession
|
||||
|
||||
spark = SparkSession.builder.config(
|
||||
"spark.jars",
|
||||
"spark-VERSION.jar", # Specify the downloaded JAR file
|
||||
)
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate()
|
||||
```
|
||||
|
||||
```scala
|
||||
import org.apache.spark.sql.SparkSession
|
||||
|
||||
val spark = SparkSession.builder
|
||||
.config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate()
|
||||
```
|
||||
|
||||
```java
|
||||
import org.apache.spark.sql.SparkSession;
|
||||
|
||||
public class QdrantSparkJavaExample {
|
||||
public static void main(String[] args) {
|
||||
SparkSession spark = SparkSession.builder()
|
||||
.config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file
|
||||
.master("local[*]")
|
||||
.appName("qdrant")
|
||||
.getOrCreate();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Loading data into Qdrant
|
||||
|
||||
<aside role="status">Before loading the data using this connector, a collection has to be <a href="/documentation/concepts/collections/#create-a-collection">created</a> in advance with the appropriate vector dimensions and configurations.</aside>
|
||||
|
||||
The connector supports ingesting multiple named/unnamed, dense/sparse vectors.
|
||||
|
||||
_Click each to expand._
|
||||
|
||||
<details>
|
||||
<summary><b>Unnamed/Default vector</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", <QDRANT_GRPC_URL>)
|
||||
.option("collection_name", <QDRANT_COLLECTION_NAME>)
|
||||
.option("embedding_field", <EMBEDDING_FIELD_NAME>) # Expected to be a field of type ArrayType(FloatType)
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Named vector</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", <QDRANT_GRPC_URL>)
|
||||
.option("collection_name", <QDRANT_COLLECTION_NAME>)
|
||||
.option("embedding_field", <EMBEDDING_FIELD_NAME>) # Expected to be a field of type ArrayType(FloatType)
|
||||
.option("vector_name", <VECTOR_NAME>)
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
> #### NOTE
|
||||
>
|
||||
> The `embedding_field` and `vector_name` options are maintained for backward compatibility. It is recommended to use `vector_fields` and `vector_names` for named vectors as shown below.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Multiple named vectors</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", "<QDRANT_GRPC_URL>")
|
||||
.option("collection_name", "<QDRANT_COLLECTION_NAME>")
|
||||
.option("vector_fields", "<COLUMN_NAME>,<ANOTHER_COLUMN_NAME>")
|
||||
.option("vector_names", "<VECTOR_NAME>,<ANOTHER_VECTOR_NAME>")
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Sparse vectors</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", "<QDRANT_GRPC_URL>")
|
||||
.option("collection_name", "<QDRANT_COLLECTION_NAME>")
|
||||
.option("sparse_vector_value_fields", "<COLUMN_NAME>")
|
||||
.option("sparse_vector_index_fields", "<COLUMN_NAME>")
|
||||
.option("sparse_vector_names", "<SPARSE_VECTOR_NAME>")
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Multiple sparse vectors</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", "<QDRANT_GRPC_URL>")
|
||||
.option("collection_name", "<QDRANT_COLLECTION_NAME>")
|
||||
.option("sparse_vector_value_fields", "<COLUMN_NAME>,<ANOTHER_COLUMN_NAME>")
|
||||
.option("sparse_vector_index_fields", "<COLUMN_NAME>,<ANOTHER_COLUMN_NAME>")
|
||||
.option("sparse_vector_names", "<SPARSE_VECTOR_NAME>,<ANOTHER_SPARSE_VECTOR_NAME>")
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Combination of named dense and sparse vectors</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", "<QDRANT_GRPC_URL>")
|
||||
.option("collection_name", "<QDRANT_COLLECTION_NAME>")
|
||||
.option("vector_fields", "<COLUMN_NAME>,<ANOTHER_COLUMN_NAME>")
|
||||
.option("vector_names", "<VECTOR_NAME>,<ANOTHER_VECTOR_NAME>")
|
||||
.option("sparse_vector_value_fields", "<COLUMN_NAME>,<ANOTHER_COLUMN_NAME>")
|
||||
.option("sparse_vector_index_fields", "<COLUMN_NAME>,<ANOTHER_COLUMN_NAME>")
|
||||
.option("sparse_vector_names", "<SPARSE_VECTOR_NAME>,<ANOTHER_SPARSE_VECTOR_NAME>")
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>No vectors - Entire dataframe is stored as payload</b></summary>
|
||||
|
||||
```python
|
||||
<pyspark.sql.DataFrame>
|
||||
.write
|
||||
.format("io.qdrant.spark.Qdrant")
|
||||
.option("qdrant_url", "<QDRANT_GRPC_URL>")
|
||||
.option("collection_name", "<QDRANT_COLLECTION_NAME>")
|
||||
.option("schema", <pyspark.sql.DataFrame>.schema.json())
|
||||
.mode("append")
|
||||
.save()
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## Databricks
|
||||
|
||||
<aside role="status">
|
||||
<p>Check out our <a href="/documentation/send-data/databricks/" target="_blank">example</a> of using the Spark connector with Databricks.</p>
|
||||
</aside>
|
||||
|
||||
You can use the `qdrant-spark` connector as a library in [Databricks](https://www.databricks.com/).
|
||||
|
||||
- Go to the `Libraries` section in your Databricks cluster dashboard.
|
||||
- Select `Install New` to open the library installation modal.
|
||||
- Search for `io.qdrant:spark:VERSION` in the Maven packages and click `Install`.
|
||||
|
||||

|
||||
|
||||
## Datatype Support
|
||||
|
||||
Qdrant supports all the Spark data types, and the appropriate data types are mapped based on the provided schema.
|
||||
|
||||
## Configuration Options
|
||||
|
||||
| Option | Description | Column DataType | Required |
|
||||
| :--------------------------- | :------------------------------------------------------------------ | :---------------------------- | :------- |
|
||||
| `qdrant_url` | GRPC URL of the Qdrant instance. Eg: <http://localhost:6334> | - | ✅ |
|
||||
| `collection_name` | Name of the collection to write data into | - | ✅ |
|
||||
| `schema` | JSON string of the dataframe schema | - | ✅ |
|
||||
| `embedding_field` | Name of the column holding the embeddings | `ArrayType(FloatType)` | ❌ |
|
||||
| `id_field` | Name of the column holding the point IDs. Default: Random UUID | `StringType` or `IntegerType` | ❌ |
|
||||
| `batch_size` | Max size of the upload batch. Default: 64 | - | ❌ |
|
||||
| `retries` | Number of upload retries. Default: 3 | - | ❌ |
|
||||
| `api_key` | Qdrant API key for authentication | - | ❌ |
|
||||
| `vector_name` | Name of the vector in the collection. | - | ❌ |
|
||||
| `vector_fields` | Comma-separated names of columns holding the vectors. | `ArrayType(FloatType)` | ❌ |
|
||||
| `vector_names` | Comma-separated names of vectors in the collection. | - | ❌ |
|
||||
| `sparse_vector_index_fields` | Comma-separated names of columns holding the sparse vector indices. | `ArrayType(IntegerType)` | ❌ |
|
||||
| `sparse_vector_value_fields` | Comma-separated names of columns holding the sparse vector values. | `ArrayType(FloatType)` | ❌ |
|
||||
| `sparse_vector_names` | Comma-separated names of the sparse vectors in the collection. | - | ❌ |
|
||||
| `shard_key_selector` | Comma-separated names of custom shard keys to use during upsert. | - | ❌ |
|
||||
|
||||
For more information, be sure to check out the [Qdrant-Spark GitHub repository](https://github.com/qdrant/qdrant-spark). The Apache Spark guide is available [here](https://spark.apache.org/docs/latest/quick-start.html). Happy data processing!
|
||||
@@ -0,0 +1,100 @@
|
||||
---
|
||||
title: Unstructured
|
||||
aliases: [ ../frameworks/unstructured/ ]
|
||||
---
|
||||
|
||||
# Unstructured
|
||||
|
||||
[Unstructured](https://unstructured.io/) is a library designed to help preprocess, structure unstructured text documents for downstream machine learning tasks.
|
||||
|
||||
Qdrant can be used as an ingestion destination in Unstructured.
|
||||
|
||||
## Setup
|
||||
|
||||
Install Unstructured with the `qdrant` extra.
|
||||
|
||||
```bash
|
||||
pip install "unstructured[qdrant]"
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
|
||||
Depending on the use case you can prefer the command line or using it within your application.
|
||||
|
||||
### CLI
|
||||
|
||||
```bash
|
||||
EMBEDDING_PROVIDER=${EMBEDDING_PROVIDER:-"langchain-huggingface"}
|
||||
|
||||
unstructured-ingest \
|
||||
local \
|
||||
--input-path example-docs/book-war-and-peace-1225p.txt \
|
||||
--output-dir local-output-to-qdrant \
|
||||
--strategy fast \
|
||||
--chunk-elements \
|
||||
--embedding-provider "$EMBEDDING_PROVIDER" \
|
||||
--num-processes 2 \
|
||||
--verbose \
|
||||
qdrant \
|
||||
--collection-name "test" \
|
||||
--url "http://localhost:6333" \
|
||||
--batch-size 80
|
||||
```
|
||||
|
||||
For a full list of the options the CLI accepts, run `unstructured-ingest <upstream connector> qdrant --help`
|
||||
|
||||
### Programmatic usage
|
||||
|
||||
```python
|
||||
from unstructured.ingest.connector.local import SimpleLocalConfig
|
||||
from unstructured.ingest.connector.qdrant import (
|
||||
QdrantWriteConfig,
|
||||
SimpleQdrantConfig,
|
||||
)
|
||||
from unstructured.ingest.interfaces import (
|
||||
ChunkingConfig,
|
||||
EmbeddingConfig,
|
||||
PartitionConfig,
|
||||
ProcessorConfig,
|
||||
ReadConfig,
|
||||
)
|
||||
from unstructured.ingest.runner import LocalRunner
|
||||
from unstructured.ingest.runner.writers.base_writer import Writer
|
||||
from unstructured.ingest.runner.writers.qdrant import QdrantWriter
|
||||
|
||||
def get_writer() -> Writer:
|
||||
return QdrantWriter(
|
||||
connector_config=SimpleQdrantConfig(
|
||||
url="http://localhost:6333",
|
||||
collection_name="test",
|
||||
),
|
||||
write_config=QdrantWriteConfig(batch_size=80),
|
||||
)
|
||||
|
||||
if __name__ == "__main__":
|
||||
writer = get_writer()
|
||||
runner = LocalRunner(
|
||||
processor_config=ProcessorConfig(
|
||||
verbose=True,
|
||||
output_dir="local-output-to-qdrant",
|
||||
num_processes=2,
|
||||
),
|
||||
connector_config=SimpleLocalConfig(
|
||||
input_path="example-docs/book-war-and-peace-1225p.txt",
|
||||
),
|
||||
read_config=ReadConfig(),
|
||||
partition_config=PartitionConfig(),
|
||||
chunking_config=ChunkingConfig(chunk_elements=True),
|
||||
embedding_config=EmbeddingConfig(provider="langchain-huggingface"),
|
||||
writer=writer,
|
||||
writer_kwargs={},
|
||||
)
|
||||
runner.run()
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- Unstructured API [reference](https://unstructured-io.github.io/unstructured/api.html).
|
||||
- Qdrant ingestion destination [reference](https://unstructured-io.github.io/unstructured/ingest/destination_connectors/qdrant.html).
|
||||
- [Source Code](https://github.com/Unstructured-IO/unstructured/blob/main/unstructured/ingest/connector/qdrant.py)
|
||||
Reference in New Issue
Block a user