Add Gemini Embedding Model 001 (#457)

* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant

* * refactor(integrations): move cohere.md to embeddings folder
* refactor(integrations): move openai.md to embeddings folder
* refactor(integrations): move autogen.md to frameworks folder
* refactor(integrations): move langchain.md to frameworks folder

* blacken

* * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both

* * chore(gemini.md): remove old Gemini integration documentation

* * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false
* chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false

* Split integrations into embedding and frameworks

* Update heading level for embedding a document

* Update Gemini embedding documentation

* Update titles for embedding and frameworks sections

* Try again with nesting

* Add documentation for integrated frameworks and embedding options

* Delete integrations documentation file

* Add Delimiter; unknown weights

* Change all weights to 3x

* Delimiter reorg

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* * docs(embedding/gemini.md): update Gemini Embedding Model API documentation
*
* - Add information about the new Gemini Embedding Model and its compatibility with Qdrant
* - Clarify the usage of the `task_type` parameter in the API call
* - Provide a list of supported task types and

* * docs(embedding): update list of embedding integrations

* * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md
* refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md

* * chore(embedding): update is_empty value to true in _index.md
* chore(embedding): remove Fifty One from embedding/_index.md

* embedding -> embeddings

---------

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
This commit is contained in:
Nirant
2023-12-11 17:50:42 +05:30
committed by GitHub
co-authored by Atita Arora
parent 6073405006
commit d90efb359d
28 changed files with 148 additions and 29 deletions
@@ -0,0 +1,73 @@
---
title: ML6 Fondant
weight: 1700
---
# ML6 Fondant
[Fondant](https://fondant.ai/en/stable/) is an open-source framework that aims to simplify and speed up large-scale data processing by making containerized components reusable across pipelines and execution environments.
Fondant features a Qdrant component image to load textual data and embeddings into a the database.
## Usage
<aside role="status">
A Qdrant collection has to be <a href="documentation/concepts/collections/">created in advance</a>
</aside>
**A data load pipeline for RAG using Qdrant**.
```python
from fondant.pipeline import ComponentOp, Pipeline
pipeline = Pipeline(
pipeline_name="ingestion-pipeline",
pipeline_description="Pipeline to prepare and process \
data for building a RAG solution",
base_path="./data-dir",
)
# An example data source component
load_from_source = ComponentOp(
component_dir="path/to/data-source-component",
arguments={
"n_rows_to_load": 10,
# Custom arguments for the component
},
)
chunk_text_op = ComponentOp.from_registry(
name="chunk_text",
arguments={
"chunk_size": 512,
"chunk_overlap": 32,
},
)
embed_text_op = ComponentOp.from_registry(
name="embed_text",
arguments={
"model_provider": "huggingface",
"model": "all-MiniLM-L6-v2",
},
)
# Getting the Qdrant component from the Fondant registry
index_qdrant_op = ComponentOp.from_registry(
name="index_qdrant",
arguments={
"url": "http:localhost:6333",
"collection_name": "some-collection-name",
},
)
# Construct your pipeline
pipeline.add_op(load_from_source)
pipeline.add_op(chunk_text_op, dependencies=load_from_source)
pipeline.add_op(embed_text_op, dependencies=chunk_text_op)
pipeline.add_op(index_qdrant_op, dependencies=embed_text_op)
```
## Next steps
Find the FondantAI docs [here](https://fondant.ai/en/stable/).