* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant * * refactor(integrations): move cohere.md to embeddings folder * refactor(integrations): move openai.md to embeddings folder * refactor(integrations): move autogen.md to frameworks folder * refactor(integrations): move langchain.md to frameworks folder * blacken * * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both * * chore(gemini.md): remove old Gemini integration documentation * * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false * chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false * Split integrations into embedding and frameworks * Update heading level for embedding a document * Update Gemini embedding documentation * Update titles for embedding and frameworks sections * Try again with nesting * Add documentation for integrated frameworks and embedding options * Delete integrations documentation file * Add Delimiter; unknown weights * Change all weights to 3x * Delimiter reorg * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * * docs(embedding/gemini.md): update Gemini Embedding Model API documentation * * - Add information about the new Gemini Embedding Model and its compatibility with Qdrant * - Clarify the usage of the `task_type` parameter in the API call * - Provide a list of supported task types and * * docs(embedding): update list of embedding integrations * * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md * refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md * * chore(embedding): update is_empty value to true in _index.md * chore(embedding): remove Fifty One from embedding/_index.md * embedding -> embeddings --------- Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
3.1 KiB
title, weight
| title | weight |
|---|---|
| Gemini | 700 |
Gemini
Qdrant is compatible with Gemini Embedding Model API and its official Python SDK that can be installed as any other package:
Gemini is a new family of Google PaLM models, released in December 2023. The new embedding models succeed the previous Gecko Embedding Model.
In the latest models, an additional parameter, task_type, can be passed to the API call. This parameter serves to designate the intended purpose for the embeddings utilized.
The Embedding Model API supports various task types, outlined as follows:
retrieval_query: Specifies the given text is a query in a search/retrieval setting.retrieval_document: Specifies the given text is a document from the corpus being searched.semantic_similarity: Specifies the given text will be used for Semantic Text Similarity.classification: Specifies that the given text will be classified.clustering: Specifies that the embeddings will be used for clustering.task_type_unspecified: Unset value, which will default to one of the other values.
If you're building a semantic search application, such as RAG, you should use task_type="retrieval_document" for the indexed documents and task_type="retrieval_query" for the search queries.
The following example shows how to do this with Qdrant:
Setup
pip install google-generativeai
Let's see how to use the Embedding Model API to embed a document for retrieval.
The following example shows how to embed a document with the models/embedding-001 with the retrieval_document task type:
Embedding a document
import pathlib
import google.generativeai as genai
import qdrant_client
GEMINI_API_KEY = "YOUR GEMINI API KEY" # add your key here
genai.configure(api_key=GEMINI_API_KEY)
result = genai.embed_content(
model="models/embedding-001",
content="Qdrant is the best vector search engine to use with Gemini",
task_type="retrieval_document",
title="Qdrant x Gemini",
)
The returned result is a dictionary with a key: embedding. The value of this key is a list of floats representing the embedding of the document.
Indexing documents with Qdrant
from qdrant_client.http.models import Batch
qdrant_client = qdrant_client.QdrantClient()
qdrant_client.upsert(
collection_name="GeminiCollection",
points=Batch(
ids=[1],
vectors=genai.embed_content(
model="models/embedding-001",
content="Qdrant is the best vector search engine to use with Gemini",
task_type="retrieval_document",
title="Qdrant x Gemini",
)["embedding"],
),
)
Searching for documents with Qdrant
Once the documents are indexed, you can search for the most relevant documents using the same model with the retrieval_query task type:
qdrant_client.search(
collection_name="GeminiCollection",
query=genai.embed_content(
model="models/embedding-001",
content="What is the best vector database to use with Gemini?",
task_type="retrieval_query",
)["embedding"],
)
That's it! You can now use Gemini Embedding Models with Qdrant.