Files
landing_page/qdrant-landing/content/documentation/embeddings/gemini.md
T
NirantandAtita Arora d90efb359d Add Gemini Embedding Model 001 (#457)
* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant

* * refactor(integrations): move cohere.md to embeddings folder
* refactor(integrations): move openai.md to embeddings folder
* refactor(integrations): move autogen.md to frameworks folder
* refactor(integrations): move langchain.md to frameworks folder

* blacken

* * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both

* * chore(gemini.md): remove old Gemini integration documentation

* * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false
* chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false

* Split integrations into embedding and frameworks

* Update heading level for embedding a document

* Update Gemini embedding documentation

* Update titles for embedding and frameworks sections

* Try again with nesting

* Add documentation for integrated frameworks and embedding options

* Delete integrations documentation file

* Add Delimiter; unknown weights

* Change all weights to 3x

* Delimiter reorg

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* Update qdrant-landing/content/documentation/embedding/gemini.md

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>

* * docs(embedding/gemini.md): update Gemini Embedding Model API documentation
*
* - Add information about the new Gemini Embedding Model and its compatibility with Qdrant
* - Clarify the usage of the `task_type` parameter in the API call
* - Provide a list of supported task types and

* * docs(embedding): update list of embedding integrations

* * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md
* refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md

* * chore(embedding): update is_empty value to true in _index.md
* chore(embedding): remove Fifty One from embedding/_index.md

* embedding -> embeddings

---------

Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
2023-12-11 17:50:42 +05:30

3.1 KiB

title, weight
title weight
Gemini 700

Gemini

Qdrant is compatible with Gemini Embedding Model API and its official Python SDK that can be installed as any other package:

Gemini is a new family of Google PaLM models, released in December 2023. The new embedding models succeed the previous Gecko Embedding Model.

In the latest models, an additional parameter, task_type, can be passed to the API call. This parameter serves to designate the intended purpose for the embeddings utilized.

The Embedding Model API supports various task types, outlined as follows:

  1. retrieval_query: Specifies the given text is a query in a search/retrieval setting.
  2. retrieval_document: Specifies the given text is a document from the corpus being searched.
  3. semantic_similarity: Specifies the given text will be used for Semantic Text Similarity.
  4. classification: Specifies that the given text will be classified.
  5. clustering: Specifies that the embeddings will be used for clustering.
  6. task_type_unspecified: Unset value, which will default to one of the other values.

If you're building a semantic search application, such as RAG, you should use task_type="retrieval_document" for the indexed documents and task_type="retrieval_query" for the search queries.

The following example shows how to do this with Qdrant:

Setup

pip install google-generativeai

Let's see how to use the Embedding Model API to embed a document for retrieval.

The following example shows how to embed a document with the models/embedding-001 with the retrieval_document task type:

Embedding a document

import pathlib
import google.generativeai as genai
import qdrant_client

GEMINI_API_KEY = "YOUR GEMINI API KEY"  # add your key here

genai.configure(api_key=GEMINI_API_KEY)

result = genai.embed_content(
    model="models/embedding-001",
    content="Qdrant is the best vector search engine to use with Gemini",
    task_type="retrieval_document",
    title="Qdrant x Gemini",
)

The returned result is a dictionary with a key: embedding. The value of this key is a list of floats representing the embedding of the document.

Indexing documents with Qdrant

from qdrant_client.http.models import Batch

qdrant_client = qdrant_client.QdrantClient()
qdrant_client.upsert(
    collection_name="GeminiCollection",
    points=Batch(
        ids=[1],
        vectors=genai.embed_content(
            model="models/embedding-001",
            content="Qdrant is the best vector search engine to use with Gemini",
            task_type="retrieval_document",
            title="Qdrant x Gemini",
        )["embedding"],
    ),
)

Searching for documents with Qdrant

Once the documents are indexed, you can search for the most relevant documents using the same model with the retrieval_query task type:

qdrant_client.search(
    collection_name="GeminiCollection",
    query=genai.embed_content(
        model="models/embedding-001",
        content="What is the best vector database to use with Gemini?",
        task_type="retrieval_query",
    )["embedding"],
)

That's it! You can now use Gemini Embedding Models with Qdrant.