add embeddings

This commit is contained in:
davidmyriel
2024-08-21 17:20:51 -07:00
parent 789cbe623c
commit 2a515dcb6f
24 changed files with 642 additions and 26 deletions
@@ -1,3 +1,4 @@
---
title: Embeddings
weight: 15
@@ -14,18 +15,30 @@ Additionally, [any open-source embeddings from HuggingFace](https://huggingface.
## Integration code samples:
| Embeddings Providers |
| ----------------------------- |
| [Aleph Alpha](./aleph-alpha/) |
| [Bedrock](./bedrock/) |
| [Cohere](./cohere/) |
| [Gemini](./gemini/) |
| [Jina](./jina-emebddngs/) |
| [Mistral](./mistral/) |
| [Nomic](./nomic/) |
| [Nvidia](./nvidia/) |
| [OpenAI](./openai/) |
| [Prem AI](./premai/) |
| [Snowflake](./snowflake/) |
| [Upstage](./upstage/) |
| [Voyage AI](./voyage/) |
| Embeddings Providers | Description |
| ----------------------------- | ----------- |
| [Aleph Alpha](./aleph-alpha/) | Multilingual embeddings focused on European languages. |
| [Bedrock](./bedrock/) | AWS managed service for foundation models and embeddings. |
| [BGE](./bge/) | Chinese embeddings for various NLP tasks. |
| [Clarifai](./clarifai/) | Embeddings for image and video recognition. |
| [Clip](./clip/) | Aligns images and text, created by OpenAI. |
| [Cohere](./cohere/) | Language model embeddings for NLP tasks. |
| [Databricks](./databricks/) | Scalable embeddings integrated with Apache Spark. |
| [Gemini](./gemini/) | Google’s multimodal embeddings for text and vision. |
| [GPT4All](./gpt4all/) | Open-source, local embeddings for privacy-focused use. |
| [Instruct](./instruct/) | Embeddings tuned for following instructions. |
| [Jina](./jina-emebddngs/) | Customizable embeddings for neural search. |
| [John Snow Labs](./johnsnow/) | Medical and clinical embeddings. |
| [Mistral](./mistral/) | Open-source, efficient language model embeddings. |
| [MixedBread](./mixedbread/) | Lightweight embeddings for constrained environments. |
| [Nomic](./nomic/) | Embeddings for data visualization. |
| [Nvidia](./nvidia_nemo/) | GPU-optimized embeddings from Nvidia. |
| [OCI](./oci/) | Oracle Cloud’s AI service with embeddings. |
| [Ollama](./ollama/) | Embeddings for conversational AI. |
| [OpenAI](./openai/) | Industry-leading embeddings for NLP. |
| [Prem AI](./premai/) | Precise language embeddings. |
| [Snowflake](./snowflake/) | Scalable embeddings for big data. |
| [Together AI](./together_ai/) | Community-driven, open-source embeddings. |
| [Upstage](./upstage/) | Embeddings for speech and language tasks. |
| [Voyage AI](./voyage/) | Navigation and spatial understanding embeddings. |
| [Watsonx](./watsonx/) | IBM's enterprise-grade embeddings. |
@@ -0,0 +1,52 @@
---
title: Clarifai
weight: 1200
aliases:
- /documentation/examples/clarifai-search/
- /documentation/tutorials/clarifai-search/
- /documentation/integrations/clarifai/
---
# Using Clarifai Embeddings with Qdrant
Clarifai is a leading provider of visual embeddings, which are particularly strong in image and video analysis. Clarifai offers an API that allows you to create embeddings for various media types, which can be integrated into Qdrant for efficient vector search and retrieval.
You can install the Clarifai Python client with pip:
```bash
pip install clarifai-client
```
## Integration Example
```python
import qdrant_client
from qdrant_client.models import Batch
from clarifai.rest import ClarifaiApp
# Initialize Clarifai client
clarifai_app = ClarifaiApp(api_key="<< your_api_key >>")
# Choose the model for embeddings
model = clarifai_app.public_models.general_embedding_model
# Upload and get embeddings for an image
image_path = "./path/to/the/image.jpg"
response = model.predict_by_filename(image_path)
# Extract the embedding from the response
embedding = response['outputs'][0]['data']['embeddings'][0]['vector']
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient()
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="MyCollection",
points=Batch(
ids=[1],
vectors=[embedding],
)
)
```
@@ -0,0 +1,55 @@
---
title: Clip
weight: 1300
aliases:
- /documentation/examples/clip-search/
- /documentation/tutorials/clip-search/
- /documentation/integrations/clip/
---
# Using Clip with Qdrant
CLIP (Contrastive Language-Image Pre-Training) provides advanced AI capabilities including natural language processing and computer vision. CLIP is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3.
## Installation
You can install the required package using the following pip command:
```bash
pip install clip-client
```
## Integration Example
```python
import qdrant_client
from qdrant_client.models import Batch
from transformers import CLIPProcessor, CLIPModel
from PIL import Image
# Load the CLIP model and processor
model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
# Load and process the image
image = Image.open("path/to/image.jpg")
inputs = processor(images=image, return_tensors="pt")
# Generate embeddings
with torch.no_grad():
embeddings = model.get_image_features(**inputs).numpy().tolist()
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="ImageEmbeddings",
points=Batch(
ids=[1],
vectors=embeddings,
)
)
```
@@ -1,6 +1,6 @@
---
title: Cohere
weight: 700
weight: 1400
aliases: [ ../integrations/cohere/ ]
---
@@ -0,0 +1,42 @@
---
title: Databricks Embeddings
weight: 1500
aliases:
- /documentation/examples/databricks-embeddings-search/
- /documentation/tutorials/databricks-embeddings-search/
- /documentation/integrations/databricks-embeddings/
---
# Using Databricks Embeddings with Qdrant
Databricks offers an advanced platform for generating embeddings, especially within large-scale data environments. You can use the following Python code to integrate Databricks-generated embeddings with Qdrant.
```python
import qdrant_client
from qdrant_client.models import Batch
from databricks import sql
# Connect to Databricks SQL endpoint
connection = sql.connect(server_hostname='your_hostname',
http_path='your_http_path',
access_token='your_access_token')
# Execute a query to get embeddings
query = "SELECT embedding FROM your_table WHERE id = 1"
cursor = connection.cursor()
cursor.execute(query)
embedding = cursor.fetchone()[0]
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="DatabricksEmbeddings",
points=Batch(
ids=[1], # Unique ID for the data point
vectors=[embedding], # Embedding fetched from Databricks
)
)
```
@@ -1,6 +1,6 @@
---
title: Gemini
weight: 700
weight: 1600
---
| Time: 10 min | Level: Beginner | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://githubtocolab.com/qdrant/examples/blob/gemini-getting-started/gemini-getting-started/gemini-getting-started.ipynb) |
@@ -0,0 +1,52 @@
---
title: GPT4All
weight: 1700
aliases:
- /documentation/examples/gpt4all-search/
- /documentation/tutorials/gpt4all-search/
- /documentation/integrations/gpt4all/
---
# Using GPT4All with Qdrant
GPT4All offers a range of large language models that can be fine-tuned for various applications. GPT4All runs large language models (LLMs) privately on everyday desktops & laptops.
No API calls or GPUs required - you can just download the application and get started. Use GPT4All in Python to program with LLMs implemented with the llama.cpp backend and Nomic's C backend.
## Installation
You can install the required package using the following pip command:
```bash
pip install gpt4all
```
Here is how you might connect to GPT4ALL using Qdrant:
```python
import qdrant_client
from qdrant_client.models import Batch
from gpt4all import GPT4All
# Initialize GPT4All model
model = GPT4All("gpt4all-lora-quantized")
# Generate embeddings for a text
text = "GPT4All enables open-source AI applications."
embeddings = model.embed(text)
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="OpenSourceAI",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -0,0 +1,47 @@
---
title: Instruct
weight: 1800
aliases:
- /documentation/examples/instruct-search/
- /documentation/tutorials/instruct-search/
- /documentation/integrations/instruct/
---
# Using Instruct with Qdrant
Instruct is a specialized provider offering detailed embeddings for instructional content, which can be effectively used with Qdrant. With Instruct every text input is embedded together with instructions explaining the use case (e.g., task and domain descriptions). Unlike encoders from prior work that are more specialized, INSTRUCTOR is a single embedder that can generate text embeddings tailored to different downstream tasks and domains, without any further training.
## Installation
```bash
pip install instruct
```
Below is an example of how to obtain embeddings using Instruct's API and store them in a Qdrant collection:
```python
import qdrant_client
from qdrant_client.models import Batch
from instruct import Instruct
# Initialize Instruct model
model = Instruct("instruct-base")
# Generate embeddings for instructional content
text = "Instruct provides detailed embeddings for learning content."
embeddings = model.embed(text)
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="LearningContent",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -1,6 +1,6 @@
---
title: Jina Embeddings
weight: 800
weight: 1900
aliases:
- /documentation/embeddings/jina-emebddngs/
- ../integrations/jina-embeddings/
@@ -0,0 +1,54 @@
---
title: John Snow Labs
weight: 2000
aliases:
- /documentation/examples/john-snow-labs-search/
- /documentation/tutorials/john-snow-labs-search/
- /documentation/integrations/john-snow-labs/
---
# Using John Snow Labs with Qdrant
John Snow Labs offers a variety of models, particularly in the healthcare domain. They have pre-trained models that can generate embeddings for medical text data.
## Installation
You can install the required package using the following pip command:
```bash
pip install johnsnowlabs
```
Here is an example of how you mmight obtain embeddings using John Snow Labs's API and store them in a Qdrant collection:
```python
import qdrant_client
from qdrant_client.models import Batch
from johnsnowlabs import nlp
# Load the pre-trained model, for example, a named entity recognition (NER) model
model = nlp.load_model("ner_jsl")
# Sample text to generate embeddings
text = "John Snow Labs provides state-of-the-art healthcare NLP solutions."
# Generate embeddings for the text
document = nlp.DocumentAssembler().setInput(text)
embeddings = model.transform(document).collectEmbeddings()
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embeddings into Qdrant
qdrant_client.upsert(
collection_name="HealthcareNLP",
points=Batch(
ids=[1], # This would be your unique ID for the data point
vectors=[embeddings],
)
)
```
@@ -1,6 +1,6 @@
---
title: Mistral
weight: 700
weight: 2100
---
| Time: 10 min | Level: Beginner | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://githubtocolab.com/qdrant/examples/blob/mistral-getting-started/mistral-embed-getting-started/mistral_qdrant_getting_started.ipynb) |
@@ -0,0 +1,51 @@
---
title: MixedBread
weight: 2200
aliases:
- /documentation/examples/mixedbread-search/
- /documentation/tutorials/mixedbread-search/
- /documentation/integrations/mixedbread/
---
# Using MixedBread with Qdrant
MixedBread is a unique provider offering embeddings across multiple domains. Their models are versatile for various search tasks when integrated with Qdrant. MixedBread is creating state-of-the-art models and tools that make search smarter, faster, and more relevant. Whether you're building a next-gen search engine or RAG (Retrieval Augmented Generation) systems, or whether you're enhancing your existing search solution, they've got the ingredients to make it happen.
## Installation
You can install the required package using the following pip command:
```bash
pip install mixedbread
```
## Integration Example
Below is an example of how to obtain embeddings using MixedBread's API and store them in a Qdrant collection:
```python
import qdrant_client
from qdrant_client.models import Batch
from mixedbread import MixedBreadModel
# Initialize MixedBread model
model = MixedBreadModel("mixedbread-variant")
# Generate embeddings
text = "MixedBread provides versatile embeddings for various domains."
embeddings = model.embed(text)
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="VersatileEmbeddings",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -1,6 +1,6 @@
---
title: "Nomic"
weight: 1100
weight: 2300
---
# Nomic
@@ -1,6 +1,6 @@
---
title: Nvidia
weight: 1200
weight: 2400
---
# Nvidia
@@ -0,0 +1,54 @@
---
title: OCI (Oracle Cloud Infrastructure)
weight: 2500
aliases:
- /documentation/examples/oci-(oracle-cloud-infrastructure)-search/
- /documentation/tutorials/oci-(oracle-cloud-infrastructure)-search/
- /documentation/integrations/oci-(oracle-cloud-infrastructure)/
---
# Using OCI (Oracle Cloud Infrastructure) with Qdrant
OCI provides robust cloud-based embeddings for various media types. The Generative AI Embedding Models convert textual input - ranging from phrases and sentences to entire paragraphs - into a structured format known as embeddings. Each piece of text input is transformed into a numerical array consisting of 1024 distinct numbers.
## Installation
You can install the required package using the following pip command:
```bash
pip install oci
```
## Code Example
Below is an example of how to obtain embeddings using OCI (Oracle Cloud Infrastructure)'s API and store them in a Qdrant collection:
```python
import qdrant_client
from qdrant_client.models import Batch
import oci
# Initialize OCI client
config = oci.config.from_file()
ai_client = oci.ai_language.AIServiceLanguageClient(config)
# Generate embeddings using OCI's AI service
text = "OCI provides cloud-based AI services."
response = ai_client.batch_detect_language_entities(text)
embeddings = response.data[0].entities[0].embedding
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="CloudAI",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -0,0 +1,52 @@
---
title: Ollama
weight: 2600
aliases:
- /documentation/examples/ollama-search/
- /documentation/tutorials/ollama-search/
- /documentation/integrations/ollama/
---
# Using Ollama with Qdrant
Ollama provides specialized embeddings for niche applications. Ollama supports a variety of embedding models, making it possible to build retrieval augmented generation (RAG) applications that combine text prompts with existing documents or other data in specialized areas.
## Installation
You can install the required package using the following pip command:
```bash
pip install ollama
```
## Integration Example
```python
import qdrant_client
from qdrant_client.models import Batch
from ollama import Ollama
# Initialize Ollama model
model = Ollama("ollama-unique")
# Generate embeddings for niche applications
text = "Ollama excels in niche applications with specific embeddings."
embeddings = model.embed(text)
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="NicheApplications",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -1,6 +1,6 @@
---
title: OpenAI
weight: 800
weight: 2700
aliases: [ ../integrations/openai/ ]
---
@@ -0,0 +1,46 @@
---
title: OpenCLIP
weight: 2750
aliases:
- /documentation/examples/openclip-search/
- /documentation/tutorials/openclip-search/
- /documentation/integrations/openclip/
---
# Using OpenCLIP with Qdrant
OpenCLIP is an open-source implementation of the CLIP model, allowing for open source generation of multimodal embeddings that link text and images.
```python
import qdrant_client
from qdrant_client.models import Batch
import open_clip
# Load the OpenCLIP model and tokenizer
model, preprocess = open_clip.create_model_and_transforms('ViT-B-32', pretrained='openai')
tokenizer = open_clip.get_tokenizer('ViT-B-32')
# Generate embeddings for a text
text = "A photo of a cat"
text_inputs = tokenizer([text])
with torch.no_grad():
text_features = model.encode_text(text_inputs)
# Convert tensor to a list
embeddings = text_features[0].cpu().numpy().tolist()
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="OpenCLIPEmbeddings",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -1,6 +1,6 @@
---
title: Prem AI
weight: 1600
weight: 2800
---
# Prem AI
@@ -1,6 +1,6 @@
---
title: Snowflake Models
weight: 1500
weight: 2900
---
# Snowflake
@@ -0,0 +1,48 @@
---
title: Together AI
weight: 3000
aliases:
- /documentation/examples/together-ai-search/
- /documentation/tutorials/together-ai-search/
- /documentation/integrations/together-ai/
---
# Using Together AI with Qdrant
Together AI focuses on collaborative AI embeddings that enhance multi-user search scenarios when integrated with Qdrant.
## Installation
You can install the required package using the following pip command:
```bash
pip install togetherai
```
## Integration Example
```python
import qdrant_client
from qdrant_client.models import Batch
from togetherai import TogetherAI
# Initialize Together AI model
model = TogetherAI("togetherai-collab")
# Generate embeddings for collaborative content
text = "Together AI enhances collaborative content search."
embeddings = model.embed(text)
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="CollaborativeContent",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```
@@ -1,6 +1,6 @@
---
title: Upstage
weight: 1700
weight: 3100
---
# Upstage
@@ -1,6 +1,6 @@
---
title: Voyage AI
weight: 1300
weight: 3200
---
# Voyage AI
@@ -0,0 +1,50 @@
---
title: Watsonx
weight: 3000
aliases:
- /documentation/examples/watsonx-search/
- /documentation/tutorials/watsonx-search/
- /documentation/integrations/watsonx/
---
# Using Watsonx with Qdrant
Watsonx is IBM's platform for AI embeddings, focusing on enterprise-level text and data analytics. These embeddings are suitable for high-precision vector searches in Qdrant.
## Installation
You can install the required package using the following pip command:
```bash
pip install watsonx
```
## Code Example
```python
import qdrant_client
from qdrant_client.models import Batch
from watsonx import Watsonx
# Initialize Watsonx AI model
model = Watsonx("watsonx-model")
# Generate embeddings for enterprise data
text = "Watsonx provides enterprise-level NLP solutions."
embeddings = model.embed(text)
# Initialize Qdrant client
qdrant_client = qdrant_client.QdrantClient(host="localhost", port=6333)
# Upsert the embedding into Qdrant
qdrant_client.upsert(
collection_name="EnterpriseData",
points=Batch(
ids=[1],
vectors=[embeddings],
)
)
```