From 829d395ebe09ff974b5416f8c0eb6417ce8f59fd Mon Sep 17 00:00:00 2001 From: Anush Date: Fri, 1 Mar 2024 00:35:24 +0530 Subject: [PATCH] docs: Added Nomic FastEmbed usage (#614) --- .../content/documentation/embeddings/nomic.md | 43 +++++++++++++++++-- .../documentation/frameworks/fondant.md | 2 +- 2 files changed, 41 insertions(+), 4 deletions(-) diff --git a/qdrant-landing/content/documentation/embeddings/nomic.md b/qdrant-landing/content/documentation/embeddings/nomic.md index c180cb8b7..6e181395e 100644 --- a/qdrant-landing/content/documentation/embeddings/nomic.md +++ b/qdrant-landing/content/documentation/embeddings/nomic.md @@ -8,12 +8,16 @@ weight: 1100 The `nomic-embed-text-v1` model is an open source [8192 context length](https://github.com/nomic-ai/contrastors) text encoder. While you can find it on the [Hugging Face Hub](https://huggingface.co/nomic-ai/nomic-embed-text-v1), you may find it easier to obtain them through the [Nomic Text Embeddings](https://docs.nomic.ai/reference/endpoints/nomic-embed-text). -Once installed, you can configure it with the official Python client or through direct HTTP requests. +Once installed, you can configure it with the official Python client, FastEmbed or through direct HTTP requests. You can use Nomic embeddings directly in Qdrant client calls. There is a difference in the way the embeddings -are obtained for documents and queries. The `task_type` parameter defines the embeddings that you get. +are obtained for documents and queries. + +#### Upsert using [Nomic SDK](https://github.com/nomic-ai/nomic) + +The `task_type` parameter defines the embeddings that you get. For documents, set the `task_type` to `search_document`: ```python @@ -36,6 +40,28 @@ qdrant_client.upsert( ) ``` +#### Upsert using [FastEmbed](https://github.com/qdrant/fastembed) + +```python +from fastembed import TextEmbedding +from qdrant_client import QdrantClient, models + +model = TextEmbedding("nomic-ai/nomic-embed-text-v1") + +output = model.embed(["Qdrant is the best vector database!"]) + +qdrant_client = QdrantClient() +qdrant_client.upsert( + collection_name="my-collection", + points=models.Batch( + ids=[1], + vectors=[embeddings.tolist() for embeddings in output], + ), +) +``` + +#### Search using [Nomic SDK](https://github.com/nomic-ai/nomic) + To query the collection, set the `task_type` to `search_query`: ```python @@ -47,7 +73,18 @@ output = embed.text( qdrant_client.search( collection_name="my-collection", - query=output["embeddings"][0], + query_vector=output["embeddings"][0], +) +``` + +#### Search using [FastEmbed](https://github.com/qdrant/fastembed) + +```python +output = next(model.embed("What is the best vector database?")) + +qdrant_client.search( + collection_name="my-collection", + query_vector=output.tolist(), ) ``` diff --git a/qdrant-landing/content/documentation/frameworks/fondant.md b/qdrant-landing/content/documentation/frameworks/fondant.md index e4bb2ea42..90995cdc4 100644 --- a/qdrant-landing/content/documentation/frameworks/fondant.md +++ b/qdrant-landing/content/documentation/frameworks/fondant.md @@ -18,7 +18,7 @@ pipeline, including a Qdrant component for writing embeddings to Qdrant. ## Usage **A data load pipeline for RAG using Qdrant**.