docs: Clean up integration docs (#1840)

* docs: Clean up integration docs

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: More cleanup

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: Deleted image artifacts

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: Updated redirects

Signed-off-by: Anush008 <anushshetty90@gmail.com>

---------

Signed-off-by: Anush008 <anushshetty90@gmail.com>
This commit is contained in:
Anush
2025-08-08 22:24:51 +05:30
committed by GitHub
parent 03455c7e46
commit 9968ce4842
43 changed files with 125 additions and 1209 deletions
@@ -16,8 +16,5 @@ partition: build
| [Confluent](/documentation/data-management/confluent/) | Fully-managed data streaming platform with a cloud-native Apache Kafka engine. |
| [DLT](/documentation/data-management/dlt/) | Python library to simplify data loading processes between several sources and destinations. |
| [Fluvio](/documentation/data-management/fluvio/) | Rust-based platform for high speed, real-time data processing. |
| [Fondant](/documentation/data-management/fondant/) | Framework for developing datasets, sharing reusable operations and data processing trees. |
| [MindsDB](/documentation/data-management/mindsdb/) | Platform to deploy, serve, and fine-tune models with numerous data source integrations. |
| [NiFi](/documentation/data-management/nifi/) | Data ingestion platform to manage data transfer between different sources and destination systems. |
| [Spark](/documentation/data-management/spark/) | A unified analytics engine for large-scale data processing. |
| [Unstructured](/documentation/data-management/unstructured/) | Python library with components for ingesting and pre-processing data from numerous sources. |
@@ -1,81 +0,0 @@
---
title: Fondant
aliases: [ ../integrations/fondant/, ../frameworks/fondant/ ]
---
# Fondant
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed
up large-scale data processing by making containerized components reusable across pipelines and
execution environments. Benefit from built-in features such as autoscaling, data lineage, and
pipeline caching, and deploy to (managed) platforms such as Vertex AI, Sagemaker, and Kubeflow
Pipelines.
Fondant comes with a library of reusable components that you can leverage to compose your own
pipeline, including a Qdrant component for writing embeddings to Qdrant.
## Usage
<aside role="status">
A Qdrant collection has to be <a href="/documentation/concepts/collections/">created in advance</a>
</aside>
**A data load pipeline for RAG using Qdrant**.
A simple ingestion pipeline could look like the following:
```python
import pyarrow as pa
from fondant.pipeline import Pipeline
indexing_pipeline = Pipeline(
name="ingestion-pipeline",
description="Pipeline to prepare and process data for building a RAG solution",
base_path="./fondant-artifacts",
)
# An custom implemenation of a read component.
text = indexing_pipeline.read(
"path/to/data-source-component",
arguments={
# your custom arguments
}
)
chunks = text.apply(
"chunk_text",
arguments={
"chunk_size": 512,
"chunk_overlap": 32,
},
)
embeddings = chunks.apply(
"embed_text",
arguments={
"model_provider": "huggingface",
"model": "all-MiniLM-L6-v2",
},
)
embeddings.write(
"index_qdrant",
arguments={
"url": "http:localhost:6333",
"collection_name": "some-collection-name",
},
cache=False,
)
```
Once you have a pipeline, you can easily run it using the built-in CLI. Fondant allows
you to run the pipeline in production across different clouds.
The first component is a custom read module that needs to be implemented and cannot be used off the
shelf. A detailed tutorial on how to rebuild this
pipeline [is provided on GitHub](https://github.com/ml6team/fondant-usecase-RAG/tree/main).
## Next steps
More information about creating your own pipelines and components can be found in the [Fondant
documentation](https://fondant.ai/en/stable/).
@@ -1,97 +0,0 @@
---
title: MindsDB
aliases: [ ../integrations/mindsdb/, ../frameworks/mindsdb/ ]
---
# MindsDB
[MindsDB](https://mindsdb.com) is an AI automation platform for building AI/ML powered features and applications. It works by connecting any source of data with any AI/ML model or framework and automating how real-time data flows between them.
With the MindsDB-Qdrant integration, you can now select Qdrant as a database to load into and retrieve from with semantic search and filtering.
**MindsDB allows you to easily**:
- Connect to any store of data or end-user application.
- Pass data to an AI model from any store of data or end-user application.
- Plug the output of an AI model into any store of data or end-user application.
- Fully automate these workflows to build AI-powered features and applications
## Usage
To get started with Qdrant and MindsDB, the following syntax can be used.
```sql
CREATE DATABASE qdrant_test
WITH ENGINE = "qdrant",
PARAMETERS = {
"location": ":memory:",
"collection_config": {
"size": 386,
"distance": "Cosine"
}
}
```
The available arguments for instantiating Qdrant can be found [here](https://github.com/mindsdb/mindsdb/blob/23a509cb26bacae9cc22475497b8644e3f3e23c3/mindsdb/integrations/handlers/qdrant_handler/qdrant_handler.py#L408-L468).
## Creating a new table
- Qdrant options for creating a collection can be specified as `collection_config` in the `CREATE DATABASE` parameters.
- By default, UUIDs are set as collection IDs. You can provide your own IDs under the `id` column.
```sql
CREATE TABLE qdrant_test.test_table (
SELECT embeddings,'{"source": "bbc"}' as metadata FROM mysql_demo_db.test_embeddings
);
```
## Querying the database
#### Perform a full retrieval using the following syntax.
```sql
SELECT * FROM qdrant_test.test_table
```
By default, the `LIMIT` is set to 10 and the `OFFSET` is set to 0.
#### Perform a similarity search using your embeddings
<aside role="status">Qdrant supports <a href="/documentation/concepts/indexing/#payload-index">payload indexing</a> that vastly improves retrieval efficiency with filters and is highly recommended. Please note that this feature currently cannot be configured via MindsDB and must be set up separately if needed.</aside>
```sql
SELECT * FROM qdrant_test.test_table
WHERE search_vector = (select embeddings from mysql_demo_db.test_embeddings limit 1)
```
#### Perform a search using filters
```sql
SELECT * FROM qdrant_test.test_table
WHERE `metadata.source` = 'bbc';
```
#### Delete entries using IDs
```sql
DELETE FROM qtest.test_table_6
WHERE id = 2
```
#### Delete entries using filters
```sql
DELETE * FROM qdrant_test.test_table
WHERE `metadata.source` = 'bbc';
```
#### Drop a table
```sql
DROP TABLE qdrant_test.test_table;
```
## Next steps
- You can find more information pertaining to MindsDB and its datasources [here](https://docs.mindsdb.com/).
- [Source Code](https://github.com/mindsdb/mindsdb/tree/main/mindsdb/integrations/handlers/qdrant_handler)
@@ -1,33 +0,0 @@
---
title: Apache NiFi
aliases: [ ../frameworks/nifi/ ]
---
# Apache NiFi
[NiFi](https://nifi.apache.org/) is a real-time data ingestion platform, which can transfer and manage data transfer between numerous sources and destination systems. It supports many protocols and offers a web-based user interface for developing and monitoring data flows.
NiFi supports ingesting and querying data in Qdrant via its processor modules.
## Configuration
![NiFi Qdrant configuration](/documentation/frameworks/nifi/nifi-conifg.png)
You can configure Qdrant NiFi processors with your Qdrant credentials, query/upload configurations. The processors offer 2 built-in embedding providers to encode data into vector embeddings - HuggingFace, OpenAI.
## Put Qdrant
![NiFI Put Qdrant](/documentation/frameworks/nifi/nifi-put-qdrant.png)
The `Put Qdrant` processor can ingest NiFi [FlowFile](https://nifi.apache.org/docs/nifi-docs/html/nifi-in-depth.html#intro) data into a Qdrant collection.
## Query Qdrant
![NiFI Query Qdrant](/documentation/frameworks/nifi/nifi-query-qdrant.png)
The `Query Qdrant` processor can perform a similarity search across a Qdrant collection and return a [FlowFile](https://nifi.apache.org/docs/nifi-docs/html/nifi-in-depth.html#intro) result.
## Further Reading
- [NiFi Documentation](https://nifi.apache.org/documentation/v2/).
- [Source Code](https://github.com/apache/nifi-python-extensions)
@@ -14,7 +14,7 @@ Qdrant can be used as an ingestion destination in Unstructured.
Install Unstructured with the `qdrant` extra.
```bash
pip install "unstructured[qdrant]"
pip install "unstructured-ingest[qdrant]"
```
## Usage
@@ -25,21 +25,21 @@ Depending on the use case you can prefer the command line or using it within you
### CLI
```bash
EMBEDDING_PROVIDER=${EMBEDDING_PROVIDER:-"langchain-huggingface"}
unstructured-ingest \
local \
--input-path example-docs/book-war-and-peace-1225p.txt \
--output-dir local-output-to-qdrant \
--strategy fast \
--chunk-elements \
--embedding-provider "$EMBEDDING_PROVIDER" \
--num-processes 2 \
--verbose \
qdrant \
--collection-name "test" \
--url "http://localhost:6333" \
--batch-size 80
--input-path $LOCAL_FILE_INPUT_DIR \
--chunking-strategy by_title \
--embedding-provider huggingface \
--partition-by-api \
--api-key $UNSTRUCTURED_API_KEY \
--partition-endpoint $UNSTRUCTURED_API_URL \
--additional-partition-args="{\"split_pdf_page\":\"true\", \"split_pdf_allow_failed\":\"true\", \"split_pdf_concurrency_level\": 15}" \
qdrant-cloud \
--url $QDRANT_URL \
--api-key $QDRANT_API_KEY \
--collection-name $QDRANT_COLLECTION \
--batch-size 50 \
--num-processes 1
```
For a full list of the options the CLI accepts, run `unstructured-ingest <upstream connector> qdrant --help`
@@ -47,54 +47,63 @@ For a full list of the options the CLI accepts, run `unstructured-ingest <upstre
### Programmatic usage
```python
from unstructured.ingest.connector.local import SimpleLocalConfig
from unstructured.ingest.connector.qdrant import (
QdrantWriteConfig,
SimpleQdrantConfig,
)
from unstructured.ingest.interfaces import (
ChunkingConfig,
EmbeddingConfig,
PartitionConfig,
ProcessorConfig,
ReadConfig,
)
from unstructured.ingest.runner import LocalRunner
from unstructured.ingest.runner.writers.base_writer import Writer
from unstructured.ingest.runner.writers.qdrant import QdrantWriter
import os
def get_writer() -> Writer:
return QdrantWriter(
connector_config=SimpleQdrantConfig(
url="http://localhost:6333",
collection_name="test",
),
write_config=QdrantWriteConfig(batch_size=80),
)
from unstructured_ingest.pipeline.pipeline import Pipeline
from unstructured_ingest.interfaces import ProcessorConfig
from unstructured_ingest.processes.connectors.local import (
LocalIndexerConfig,
LocalDownloaderConfig,
LocalConnectionConfig
)
from unstructured_ingest.processes.partitioner import PartitionerConfig
from unstructured_ingest.processes.chunker import ChunkerConfig
from unstructured_ingest.processes.embedder import EmbedderConfig
from unstructured_ingest.processes.connectors.qdrant.cloud import (
CloudQdrantConnectionConfig,
CloudQdrantAccessConfig,
CloudQdrantUploadStagerConfig,
CloudQdrantUploaderConfig
)
if __name__ == "__main__":
writer = get_writer()
runner = LocalRunner(
processor_config=ProcessorConfig(
verbose=True,
output_dir="local-output-to-qdrant",
num_processes=2,
Pipeline.from_configs(
context=ProcessorConfig(),
indexer_config=LocalIndexerConfig(input_path=os.getenv("LOCAL_FILE_INPUT_DIR")),
downloader_config=LocalDownloaderConfig(),
source_connection_config=LocalConnectionConfig(),
partitioner_config=PartitionerConfig(
partition_by_api=True,
api_key=os.getenv("UNSTRUCTURED_API_KEY"),
partition_endpoint=os.getenv("UNSTRUCTURED_API_URL"),
additional_partition_args={
"split_pdf_page": True,
"split_pdf_allow_failed": True,
"split_pdf_concurrency_level": 15
}
),
connector_config=SimpleLocalConfig(
input_path="example-docs/book-war-and-peace-1225p.txt",
chunker_config=ChunkerConfig(chunking_strategy="by_title"),
embedder_config=EmbedderConfig(embedding_provider="huggingface"),
destination_connection_config=CloudQdrantConnectionConfig(
access_config=CloudQdrantAccessConfig(
api_key=os.getenv("QDRANT_API_KEY")
),
url=os.getenv("QDRANT_URL")
),
read_config=ReadConfig(),
partition_config=PartitionConfig(),
chunking_config=ChunkingConfig(chunk_elements=True),
embedding_config=EmbeddingConfig(provider="langchain-huggingface"),
writer=writer,
writer_kwargs={},
)
runner.run()
stager_config=CloudQdrantUploadStagerConfig(),
uploader_config=CloudQdrantUploaderConfig(
collection_name=os.getenv("QDRANT_COLLECTION"),
batch_size=50,
num_processes=1
)
).run()
```
## Next steps
- Unstructured API [reference](https://unstructured-io.github.io/unstructured/api.html).
- Qdrant ingestion destination [reference](https://unstructured-io.github.io/unstructured/ingest/destination_connectors/qdrant.html).
- [Source Code](https://github.com/Unstructured-IO/unstructured-ingest/blob/main/unstructured_ingest/connector/qdrant.py)
- Qdrant ingestion destination [reference](https://docs.unstructured.io/ui/destinations/qdrant).
- [Source Code](https://github.com/Unstructured-IO/unstructured-ingest/tree/main/unstructured_ingest/processes/connectors/qdrant)