Update Fondant documentation (#476)

This commit is contained in:
Matthias Richter
2024-01-04 20:07:26 +05:30
committed by GitHub
parent 6b824222bb
commit 2f2a2d6695
@@ -1,14 +1,19 @@
---
title: ML6 Fondant
title: Fondant
weight: 1700
aliases: [ ../integrations/fondant/ ]
---
# ML6 Fondant
# Fondant
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed up large-scale data processing by making containerized components reusable across pipelines and execution environments.
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed
up large-scale data processing by making containerized components reusable across pipelines and
execution environments. Benefit from built-in features such as autoscaling, data lineage, and
pipeline caching, and deploy to (managed) platforms such as Vertex AI, Sagemaker, and Kubeflow
Pipelines.
Fondant features a Qdrant component image to load textual data and embeddings into a the database.
Fondant comes with a library of reusable components that you can leverage to compose your own
pipeline, including a Qdrant component for writing embeddings to Qdrant.
## Usage
@@ -18,57 +23,60 @@ A Qdrant collection has to be <a href="documentation/concepts/collections/">crea
**A data load pipeline for RAG using Qdrant**.
A simple ingestion pipeline could look like the following:
```python
from fondant.pipeline import ComponentOp, Pipeline
import pyarrow as pa
from fondant.pipeline import Pipeline
pipeline = Pipeline(
pipeline_name="ingestion-pipeline",
pipeline_description="Pipeline to prepare and process \
data for building a RAG solution",
base_path="./data-dir",
indexing_pipeline = Pipeline(
name="ingestion-pipeline",
description="Pipeline to prepare and process data for building a RAG solution",
base_path="./fondant-artifacts",
)
# An example data source component
load_from_source = ComponentOp(
component_dir="path/to/data-source-component",
# An custom implemenation of a read component.
text = indexing_pipeline.read(
"path/to/data-source-component",
arguments={
"n_rows_to_load": 10,
# Custom arguments for the component
},
# your custom arguments
}
)
chunk_text_op = ComponentOp.from_registry(
name="chunk_text",
chunks = text.apply(
"chunk_text",
arguments={
"chunk_size": 512,
"chunk_overlap": 32,
},
)
embed_text_op = ComponentOp.from_registry(
name="embed_text",
embeddings = chunks.apply(
"embed_text",
arguments={
"model_provider": "huggingface",
"model": "all-MiniLM-L6-v2",
},
)
# Getting the Qdrant component from the Fondant registry
index_qdrant_op = ComponentOp.from_registry(
name="index_qdrant",
embeddings.write(
"index_qdrant",
arguments={
"url": "http:localhost:6333",
"collection_name": "some-collection-name",
},
cache=False,
)
# Construct your pipeline
pipeline.add_op(load_from_source)
pipeline.add_op(chunk_text_op, dependencies=load_from_source)
pipeline.add_op(embed_text_op, dependencies=chunk_text_op)
pipeline.add_op(index_qdrant_op, dependencies=embed_text_op)
```
Once you have a pipeline, you can easily run it using the built-in CLI. Fondant allows
you to run the pipeline in production across different clouds.
The first component is a custom read module that needs to be implemented and cannot be used off the
shelf. A detailed tutorial on how to rebuild this
pipeline [is provided on GitHub](https://github.com/ml6team/fondant-usecase-RAG/tree/main).
## Next steps
Find the FondantAI docs [here](https://fondant.ai/en/stable/).
More information about creating your own pipelines and components can be found in the [Fondant
documentation](https://fondant.ai/en/stable/).