docs: Added Nomic FastEmbed usage (#614)

This commit is contained in:
Anush
2024-03-01 00:35:24 +05:30
committed by GitHub
parent 6a5645c001
commit 829d395ebe
2 changed files with 41 additions and 4 deletions
@@ -8,12 +8,16 @@ weight: 1100
The `nomic-embed-text-v1` model is an open source [8192 context length](https://github.com/nomic-ai/contrastors) text encoder.
While you can find it on the [Hugging Face Hub](https://huggingface.co/nomic-ai/nomic-embed-text-v1),
you may find it easier to obtain them through the [Nomic Text Embeddings](https://docs.nomic.ai/reference/endpoints/nomic-embed-text).
Once installed, you can configure it with the official Python client or through direct HTTP requests.
Once installed, you can configure it with the official Python client, FastEmbed or through direct HTTP requests.
<aside role="status">Using Nomic Text Embeddings requires configuring the Nomic API token</aside>
You can use Nomic embeddings directly in Qdrant client calls. There is a difference in the way the embeddings
are obtained for documents and queries. The `task_type` parameter defines the embeddings that you get.
are obtained for documents and queries.
#### Upsert using [Nomic SDK](https://github.com/nomic-ai/nomic)
The `task_type` parameter defines the embeddings that you get.
For documents, set the `task_type` to `search_document`:
```python
@@ -36,6 +40,28 @@ qdrant_client.upsert(
)
```
#### Upsert using [FastEmbed](https://github.com/qdrant/fastembed)
```python
from fastembed import TextEmbedding
from qdrant_client import QdrantClient, models
model = TextEmbedding("nomic-ai/nomic-embed-text-v1")
output = model.embed(["Qdrant is the best vector database!"])
qdrant_client = QdrantClient()
qdrant_client.upsert(
collection_name="my-collection",
points=models.Batch(
ids=[1],
vectors=[embeddings.tolist() for embeddings in output],
),
)
```
#### Search using [Nomic SDK](https://github.com/nomic-ai/nomic)
To query the collection, set the `task_type` to `search_query`:
```python
@@ -47,7 +73,18 @@ output = embed.text(
qdrant_client.search(
collection_name="my-collection",
query=output["embeddings"][0],
query_vector=output["embeddings"][0],
)
```
#### Search using [FastEmbed](https://github.com/qdrant/fastembed)
```python
output = next(model.embed("What is the best vector database?"))
qdrant_client.search(
collection_name="my-collection",
query_vector=output.tolist(),
)
```
@@ -18,7 +18,7 @@ pipeline, including a Qdrant component for writing embeddings to Qdrant.
## Usage
<aside role="status">
A Qdrant collection has to be <a href="documentation/concepts/collections/">created in advance</a>
A Qdrant collection has to be <a href="/documentation/concepts/collections/">created in advance</a>
</aside>
**A data load pipeline for RAG using Qdrant**.