Add Gemini Embedding Model 001 (#457)

* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant

* * refactor(integrations): move cohere.md to embeddings folder
* refactor(integrations): move openai.md to embeddings folder
* refactor(integrations): move autogen.md to frameworks folder
* refactor(integrations): move langchain.md to frameworks folder

* blacken

* * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both

* * chore(gemini.md): remove old Gemini integration documentation

* * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false
* chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false

* Split integrations into embedding and frameworks

* Update heading level for embedding a document

* Update Gemini embedding documentation

* Update titles for embedding and frameworks sections

* Try again with nesting

* Add documentation for integrated frameworks and embedding options

* Delete integrations documentation file

* Add Delimiter; unknown weights

* Change all weights to 3x

* Delimiter reorg

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* * docs(embedding/gemini.md): update Gemini Embedding Model API documentation
*
* - Add information about the new Gemini Embedding Model and its compatibility with Qdrant
* - Clarify the usage of the `task_type` parameter in the API call
* - Provide a list of supported task types and

* * docs(embedding): update list of embedding integrations

* * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md
* refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md

* * chore(embedding): update is_empty value to true in _index.md
* chore(embedding): remove Fifty One from embedding/_index.md

* embedding -> embeddings

---------

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
This commit is contained in:
Nirant
2023-12-11 17:50:42 +05:30
committed by GitHub
co-authored by Atita Arora
parent 6073405006
commit d90efb359d
28 changed files with 148 additions and 29 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
--- ---
#Delimiter files are used to separate the list of documentation pages into sections. #Delimiter files are used to separate the list of documentation pages into sections.
title: "Support" title: "Integrations"
type: delimiter type: delimiter
weight: 30 # Change this weight to change order of sections weight: 30 # Change this weight to change order of sections
sitemapExclude: True sitemapExclude: True
@@ -0,0 +1,7 @@
---
#Delimiter files are used to separate the list of documentation pages into sections.
title: "Support"
type: delimiter
weight: 40 # Change this weight to change order of sections
sitemapExclude: True
---
@@ -1,6 +1,6 @@
--- ---
title: Community links title: Community links
weight: 32 weight: 42
--- ---
# Community Contributions # Community Contributions
@@ -0,0 +1,14 @@
---
title: Embeddings
weight: 33
# If the index.md file is empty, the link to the section will be hidden from the sidebar
is_empty: true
---
| Embedding |
|---|
| [Gemini](./gemini/) |
| [Aleph Alpha](./aleph-alpha/) |
| [Cohere](./cohere/) |
| [Jina](./jina-emebddngs/) |
| [OpenAI](./openai/) |
@@ -0,0 +1,94 @@
---
title: Gemini
weight: 700
---
# Gemini
Qdrant is compatible with Gemini Embedding Model API and its official Python SDK that can be installed as any other package:
Gemini is a new family of Google PaLM models, released in December 2023. The new embedding models succeed the previous Gecko Embedding Model.
In the latest models, an additional parameter, `task_type`, can be passed to the API call. This parameter serves to designate the intended purpose for the embeddings utilized.
The Embedding Model API supports various task types, outlined as follows:
1. `retrieval_query`: Specifies the given text is a query in a search/retrieval setting.
2. `retrieval_document`: Specifies the given text is a document from the corpus being searched.
3. `semantic_similarity`: Specifies the given text will be used for Semantic Text Similarity.
4. `classification`: Specifies that the given text will be classified.
5. `clustering`: Specifies that the embeddings will be used for clustering.
6. `task_type_unspecified`: Unset value, which will default to one of the other values.
If you're building a semantic search application, such as RAG, you should use `task_type="retrieval_document"` for the indexed documents and `task_type="retrieval_query"` for the search queries.
The following example shows how to do this with Qdrant:
## Setup
```bash
pip install google-generativeai
```
Let's see how to use the Embedding Model API to embed a document for retrieval.
The following example shows how to embed a document with the `models/embedding-001` with the `retrieval_document` task type:
## Embedding a document
```python
import pathlib
import google.generativeai as genai
import qdrant_client
GEMINI_API_KEY = "YOUR GEMINI API KEY" # add your key here
genai.configure(api_key=GEMINI_API_KEY)
result = genai.embed_content(
model="models/embedding-001",
content="Qdrant is the best vector search engine to use with Gemini",
task_type="retrieval_document",
title="Qdrant x Gemini",
)
```
The returned result is a dictionary with a key: `embedding`. The value of this key is a list of floats representing the embedding of the document.
## Indexing documents with Qdrant
```python
from qdrant_client.http.models import Batch
qdrant_client = qdrant_client.QdrantClient()
qdrant_client.upsert(
collection_name="GeminiCollection",
points=Batch(
ids=[1],
vectors=genai.embed_content(
model="models/embedding-001",
content="Qdrant is the best vector search engine to use with Gemini",
task_type="retrieval_document",
title="Qdrant x Gemini",
)["embedding"],
),
)
```
## Searching for documents with Qdrant
Once the documents are indexed, you can search for the most relevant documents using the same model with the `retrieval_query` task type:
```python
qdrant_client.search(
collection_name="GeminiCollection",
query=genai.embed_content(
model="models/embedding-001",
content="What is the best vector database to use with Gemini?",
task_type="retrieval_query",
)["embedding"],
)
```
That's it! You can now use Gemini Embedding Models with Qdrant.
@@ -5,16 +5,15 @@ weight: 800
# OpenAI # OpenAI
Qdrant can also easily work with [OpenAI embeddings](https://beta.openai.com/docs/guides/embeddings/embeddings). There is an Qdrant can also easily work with [OpenAI embeddings](https://beta.openai.com/docs/guides/embeddings/embeddings).
official OpenAI Python package that simplifies obtaining them, and it might be installed with pip:
There is an official OpenAI Python package that simplifies obtaining them, and it might be installed with pip:
```bash ```bash
pip install openai pip install openai
``` ```
Once installed, the package exposes the method allowing to retrieve the embedding for given text. OpenAI requires an API key Once installed, the package exposes the method allowing to retrieve the embedding for given text. OpenAI requires an API key that has to be provided either as an environmental variable `OPENAI_API_KEY` or set in the source code directly, as presented below:
that has to be provided either as an environmental variable `OPENAI_API_KEY` or set in the source code directly, as
presented below:
```python ```python
import openai import openai
@@ -38,7 +37,7 @@ qdrant_client.upsert(
points=Batch( points=Batch(
ids=[1], ids=[1],
vectors=[response["data"][0]["embedding"]], vectors=[response["data"][0]["embedding"]],
) ),
) )
``` ```
@@ -1,5 +1,5 @@
--- ---
title: FAQ title: FAQ
weight: 31 weight: 41
is_empty: true is_empty: true
--- ---
@@ -0,0 +1,24 @@
---
title: Frameworks
weight: 33
# If the index.md file is empty, the link to the section will be hidden from the sidebar
is_empty: true
---
| Frameworks |
|---|
| [AirByte](./airbyte/) |
| [AutoGen](./autogen/) |
| [Cheshire Cat](./cheshire-cat/) |
| [DLT](./dlt/) |
| [DocArray](./docarray/) |
| [DSPy](./dspy/) |
| [Fifty One](./fifty-one/) |
| [txtai](./txtai/) |
| [Fondant](./fondant/) |
| [Haystack](./haystack/) |
| [Langchain](./langchain/) |
| [Llama Index](./llama-index/) |
| [Minds DB](./mindsdb/) |
| [PrivateGPT](./privategpt/) |
| [Spark](./spark/) |
@@ -1,19 +0,0 @@
---
title: Integrations
weight: 24
# If the index.md file is empty, the link to the section will be hidden from the sidebar
is_empty: true
---
| Integration Instructions | Examples |
|---|---|
|AlephAlpha| Multimodal Search |
|Cohere | Adaptive Q&A Engine|
|DocArray |
|FiftyOne|
|Haystack |
|LangChain |
|LlamaIndex | Adaptive Q&A Engine|
|OpenAI |
|txtai |
@@ -1,6 +1,6 @@
--- ---
title: Release notes title: Release notes
weight: 32 weight: 42
type: external-link type: external-link
external_url: https://github.com/qdrant/qdrant/releases external_url: https://github.com/qdrant/qdrant/releases
sitemapExclude: True sitemapExclude: True