mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-01 08:58:31 +02:00
@@ -5,9 +5,9 @@ weight: 81
|
||||
|
||||
# Inference in Qdrant Managed Cloud
|
||||
|
||||
Inference is the process of creating vector embeddings from text, images, or other data types using a machine learning model.
|
||||
[Inference](/documentation/concepts/inference/) is the process of creating vector embeddings from text, images, or other data types using a machine learning model.
|
||||
|
||||
Qdrant Managed Cloud allows you to use inference directly in the cloud, without the need to set up and maintain your own inference infrastructure.
|
||||
Qdrant Managed Cloud allows you to use inference directly in the cloud, without the need to set up and maintain your own inference infrastructure. You can use [embedding models hosted on Qdrant Cloud](#cloud-inference), or use [externally hosted models](#use-external-models).
|
||||
|
||||
<aside role="alert">
|
||||
Inference is currently executed within a US region, even if the Qdrant Cloud cluster is hosted in another region.
|
||||
@@ -15,130 +15,34 @@ Qdrant Managed Cloud allows you to use inference directly in the cloud, without
|
||||
|
||||

|
||||
|
||||
## Supported Models
|
||||
|
||||
You can see the list of supported models in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. The list includes models for text, both to produce dense and sparse vectors, as well as multi-modal models for images.
|
||||
|
||||
## Enabling/Disabling Inference
|
||||
|
||||
Inference is enabled by default for all new clusters, created after July, 7th 2025. You can enable it for existing clusters directly from the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. Activating inference will trigger a restart of your cluster to apply the new configuration.
|
||||
|
||||
## Billing
|
||||
|
||||
Inference is billed based on the number of tokens processed by the model. The cost is calculated per 1,000,000 tokens. The price depends on the model and is displayed ont the Inference tab of the Cluster Detail page. You also can see the current usage of each model there.
|
||||
|
||||
## Using Inference
|
||||
|
||||
Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs when upserting points and when querying the database.
|
||||
Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs when upserting points and when querying the database. Refer to the [Inference documentation](/documentation/concepts/inference/) for details.
|
||||
|
||||
Instead of a vector, you can use special *Interface Objects*:
|
||||
## Cloud Inference
|
||||
|
||||
* **`Document`** object, used for text inference
|
||||
Clusters on Qdrant Managed Cloud can access embedding models that are hosted on Qdrant Cloud.
|
||||
|
||||
```js
|
||||
// Document
|
||||
{
|
||||
// Text input
|
||||
text: "Your text",
|
||||
// Name of the model, to do inference with
|
||||
model: "<the-model-to-use>",
|
||||
// Extra parameters for the model, Optional
|
||||
options: {}
|
||||
}
|
||||
```
|
||||
You can see the list of supported models in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. The list includes models for text, both to produce dense and sparse vectors, as well as multi-modal models for images.
|
||||
|
||||
* **`Image`** object, used for image inference
|
||||
### Billing
|
||||
|
||||
```js
|
||||
// Image
|
||||
{
|
||||
// Image input
|
||||
image: "<url>", // Or base64 encoded image
|
||||
// Name of the model, to do inference with
|
||||
model: "<the-model-to-use>",
|
||||
// Extra parameters for the model, Optional
|
||||
options: {}
|
||||
}
|
||||
```
|
||||
Inference is billed based on the number of tokens processed by the model. The cost is calculated per 1,000,000 tokens. The price depends on the model and is displayed on the Inference tab of the Cluster Detail page. You also can see the current usage of each model there.
|
||||
|
||||
* **`Object`** object, reserved for other types of input, which might be implemented in the future.
|
||||
## Use External Models
|
||||
|
||||
Qdrant Cloud can act as a proxy for the APIs of three external embedding model providers:
|
||||
|
||||
The Qdrant API supports usage of these Inference Objects in all places, where regular vectors can be used.
|
||||
- OpenAI
|
||||
- Cohere
|
||||
- Jina AI
|
||||
|
||||
For example:
|
||||
This enables you to access any of the embedding models provided by these providers through the Qdrant API.
|
||||
|
||||
```http
|
||||
POST /collections/<your-collection>/points/query
|
||||
{
|
||||
"query": {
|
||||
"nearest": [0.12, 0.34, 0.56, 0.78, ...]
|
||||
}
|
||||
}
|
||||
```
|
||||
### Billing
|
||||
|
||||
Can be replaced with
|
||||
|
||||
```http
|
||||
POST /collections/<your-collection>/points/query
|
||||
{
|
||||
"query": {
|
||||
"nearest": {
|
||||
"text": "My Query Text",
|
||||
"model": "<the-model-to-use>"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
In this case, the Qdrant Cloud will use the configured embedding model to automatically create a vector from the Inference Object and then perform the search query with it. All of this happens within a low-latency network.
|
||||
|
||||
The input used for inference will not be saved anywhere. If you want to persist it in Qdrant, make sure to explicitly include it in the payload.
|
||||
|
||||
|
||||
### Text Inference
|
||||
|
||||
Let's consider an example of using Cloud Inference with a text model producing dense vectors.
|
||||
|
||||
Here, we create one point and use a simple search query with a `Document` Inference Object.
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/simple/" >}}
|
||||
|
||||
Usage examples, specific to each cluster and model, can also be found in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console.
|
||||
|
||||
Note that each model has a context window, which is the maximum number of tokens that can be processed by the model in a single request. If the input text exceeds the context window, it will be truncated to fit within the limit. The context window size is displayed in the Inference tab of the Cluster Detail page.
|
||||
|
||||
For dense vector models, you also have to ensure that the vector size configured in the collection matches the output size of the model. If the vector size does not match, the upsert will fail with an error.
|
||||
|
||||
### Image Inference
|
||||
|
||||
Here is another example of using Cloud Inference with an image model. This time, we will use the `CLIP` model to encode an image and then use a text query to search for it.
|
||||
|
||||
Since the `CLIP` model is multimodal, we can use both image and text inputs on the same vector field.
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/image/" >}}
|
||||
|
||||
Qdrant Cloud Inference server will download the images using the provided link. Alternatively, you can upload the image as a base64 encoded string.
|
||||
|
||||
Note that each model has limitations on the file size and extensions it can work with.
|
||||
|
||||
Please refer to the model card for details.
|
||||
|
||||
### Local Inference Compatibility
|
||||
|
||||
The Python SDK offers a unique capability: it supports both [local](/documentation/fastembed/fastembed-semantic-search/) and cloud inference through an identical interface.
|
||||
|
||||
You can easily switch between local and cloud inference by setting the cloud_inference flag when initializing the QdrantClient. For example:
|
||||
|
||||
```python
|
||||
client = QdrantClient(
|
||||
url="https://your-cluster.qdrant.io",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True, # Set to False to use local inference
|
||||
)
|
||||
```
|
||||
|
||||
This flexibility allows you to develop and test your applications locally or in continuous integration (CI) environments without requiring access to cloud inference resources.
|
||||
|
||||
* When `cloud_inference` is set to `False`, inference is performed locally usign `fastembed`.
|
||||
* When set to `True`, inference requests are handled by Qdrant Cloud.
|
||||
To use an external provider's embedding model, you need an API key from that provider. Billing is managed directly through the external provider, based on API key usage. Refer to each external embedding model provider's website for pricing details.
|
||||
@@ -39,6 +39,10 @@ A [Payload](/documentation/concepts/payload/) describes information that you can
|
||||
|
||||
[Filtering](/documentation/concepts/filtering/) defines various database-style clauses, conditions, and more.
|
||||
|
||||
## Inference
|
||||
|
||||
[Inference](/documentation/concepts/inference/) is the process of generating vectors (embeddings) from text or an image.
|
||||
|
||||
## Optimizer
|
||||
|
||||
[Optimizer](/documentation/concepts/optimizer/) describes options to rebuild
|
||||
|
||||
@@ -0,0 +1,248 @@
|
||||
---
|
||||
title: Inference
|
||||
weight: 65
|
||||
aliases:
|
||||
- ../inference
|
||||
---
|
||||
|
||||
# Inference
|
||||
|
||||
Inference is the process of using a machine learning model to create vector embeddings from text, images, or other data types. While you can create embeddings on the client side, you can also let Qdrant generate them while storing or querying data.
|
||||
|
||||

|
||||
|
||||
There are several advantages to generating embeddings with Qdrant:
|
||||
|
||||
- No need for external pipelines or separate model servers.
|
||||
- Work with a single unified API instead of a different API per model provider.
|
||||
- No external network calls, minimizing delays or data transfer overhead.
|
||||
|
||||
Depending on the model you want to use, inference can be executed:
|
||||
|
||||
- on the client side, using the [FastEmbed](/documentation/fastembed/) library
|
||||
- [by the Qdrant cluster](#server-side-inference-bm25) (only supported for the BM25 model)
|
||||
- in Qdrant Cloud, using [Cloud Inference](#qdrant-cloud-inference) (for clusters on Qdrant Managed Cloud)
|
||||
- [externally](#external-embedding-model-providers) (models by OpenAI, Cohere, and Jina AI; for clusters on Qdrant Managed Cloud)
|
||||
|
||||
## Inference API
|
||||
|
||||
You can use inference in the API wherever you can use regular vectors. Instead of a vector, you can use special *Interface Objects*:
|
||||
|
||||
* **`Document`** object, used for text inference
|
||||
|
||||
```js
|
||||
// Document
|
||||
{
|
||||
// Text input
|
||||
text: "Your text",
|
||||
// Name of the model, to do inference with
|
||||
model: "<the-model-to-use>",
|
||||
// Extra parameters for the model, Optional
|
||||
options: {}
|
||||
}
|
||||
```
|
||||
|
||||
* **`Image`** object, used for image inference
|
||||
|
||||
```js
|
||||
// Image
|
||||
{
|
||||
// Image input
|
||||
image: "<url>", // Or base64 encoded image
|
||||
// Name of the model, to do inference with
|
||||
model: "<the-model-to-use>",
|
||||
// Extra parameters for the model, Optional
|
||||
options: {}
|
||||
}
|
||||
```
|
||||
|
||||
* **`Object`** object, reserved for other types of input, which might be implemented in the future.
|
||||
|
||||
|
||||
The Qdrant API supports the usage of these Inference Objects in all places where regular vectors can be used. For example:
|
||||
|
||||
```http
|
||||
POST /collections/<your-collection>/points/query
|
||||
{
|
||||
"query": {
|
||||
"nearest": [0.12, 0.34, 0.56, 0.78, ...]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Can be replaced with
|
||||
|
||||
```http
|
||||
POST /collections/<your-collection>/points/query
|
||||
{
|
||||
"query": {
|
||||
"nearest": {
|
||||
"text": "My Query Text",
|
||||
"model": "<the-model-to-use>"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
In this case, Qdrant uses the configured embedding model to automatically create a vector from the Inference Object and then perform the search query with it. All of this happens within a low-latency network.
|
||||
|
||||
<aside role="status">
|
||||
When using inference at ingest time, the input used for inference is not stored. If you want to persist it in Qdrant, ensure that you explicitly include it in the payload.
|
||||
</aside>
|
||||
|
||||
## Server-side Inference: BM25
|
||||
|
||||
BM25 (Best Matching 25) is a ranking function for text search. BM25 uses sparse vectors that represent documents, where each dimension corresponds to a word. Qdrant can generate these sparse embeddings from input text directly on the server.
|
||||
|
||||
While upserting points, provide the text and the `qdrant/bm25` embedding model:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/ingest/" >}}
|
||||
|
||||
Qdrant uses the model to generate the embeddings and stores the point with the resulting vector. Retrieving the point shows the embeddings that were generated:
|
||||
|
||||
```json
|
||||
....
|
||||
"my-bm25-vector": {
|
||||
"indices": [
|
||||
112174620,
|
||||
177304315,
|
||||
662344706,
|
||||
771857363,
|
||||
1617337648
|
||||
],
|
||||
"values": [
|
||||
1.6697302,
|
||||
1.6697302,
|
||||
1.6697302,
|
||||
1.6697302,
|
||||
1.6697302
|
||||
]
|
||||
}
|
||||
....
|
||||
]
|
||||
```
|
||||
|
||||
Similarly, you can use inference at query time by providing the text to query with as well as the embedding model:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/query/" >}}
|
||||
|
||||
## Qdrant Cloud Inference
|
||||
|
||||
Clusters on Qdrant Managed Cloud can access embedding models that are [hosted on Qdrant Cloud](/documentation/cloud/inference/). For a list of available models, visit the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. Here, you can also enable Cloud Inference for a cluster if it's not already enabled.
|
||||
|
||||
Before using a Cloud-hosted embedding model, ensure that your collection has been configured for vectors with the correct dimensionality. The Inference tab of the Cluster Detail page in the Qdrant Cloud Console lists the dimensionality for each supported embedding model.
|
||||
|
||||
### Text Inference
|
||||
|
||||
Let's consider an example of using Cloud Inference with a text model that produces dense vectors. This example creates one point and uses a simple search query with a `Document` Inference Object.
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/simple/" >}}
|
||||
|
||||
Usage examples, specific to each cluster and model, can also be found in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console.
|
||||
|
||||
Note that each model has a context window, which is the maximum number of tokens that can be processed by the model in a single request. If the input text exceeds the context window, it is truncated to fit within the limit. The context window size is displayed in the Inference tab of the Cluster Detail page.
|
||||
|
||||
For dense vector models, you also have to ensure that the vector size configured in the collection matches the output size of the model. If the vector size does not match, the upsert will fail with an error.
|
||||
|
||||
### Image Inference
|
||||
|
||||
Here is another example of using Cloud Inference with an image model. This example uses the `CLIP` model to encode an image and then uses a text query to search for it.
|
||||
|
||||
Since the `CLIP` model is multimodal, we can use both image and text inputs on the same vector field.
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/image/" >}}
|
||||
|
||||
The Qdrant Cloud Inference server will download the images using the provided URL. Alternatively, you can provide the image as a base64-encoded string. Each model has limitations on the file size and extensions it can work with. Refer to the model card for details.
|
||||
|
||||
### Local Inference Compatibility
|
||||
|
||||
The Python SDK offers a unique capability: it supports both [local](/documentation/fastembed/fastembed-semantic-search/) and cloud inference through an identical interface.
|
||||
|
||||
You can easily switch between local and cloud inference by setting the `cloud_inference` flag when initializing the QdrantClient. For example:
|
||||
|
||||
```python
|
||||
client = QdrantClient(
|
||||
url="https://your-cluster.qdrant.io",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True, # Set to False to use local inference
|
||||
)
|
||||
```
|
||||
|
||||
This flexibility allows you to develop and test your applications locally or in continuous integration (CI) environments without requiring access to cloud inference resources.
|
||||
|
||||
* When `cloud_inference` is set to `False`, inference is performed locally using `fastembed`.
|
||||
* When set to `True`, inference requests are handled by Qdrant Cloud.
|
||||
|
||||
## External Embedding Model Providers
|
||||
|
||||
Qdrant Cloud can act as a proxy for the APIs of three external embedding model providers:
|
||||
|
||||
- OpenAI
|
||||
- Cohere
|
||||
- Jina AI
|
||||
|
||||
This enables you to access any of the embedding models provided by these providers through the Qdrant API.
|
||||
|
||||
To use an external provider's embedding model, you need an API key from that provider. For example, to access OpenAI models, you need an OpenAI API key. Qdrant does not store or cache your API keys; they must be provided with each inference request.
|
||||
|
||||
When using an external embedding model, ensure that your collection has been configured for vectors with the correct dimensionality. Refer to the model's documentation for details on the output dimensions.
|
||||
|
||||
<aside role="status">
|
||||
When using a model from an external provider, refer to the model's documentation for:
|
||||
|
||||
- the dimensions of the resulting embeddings
|
||||
- how to pass an image when creating image embeddings. Some providers allow you to pass an image URL, while others require a base64-encoded image
|
||||
- any additional parameters that the model supports
|
||||
</aside>
|
||||
|
||||
### OpenAI
|
||||
|
||||
When you prepend a model name with `openai/`, the embedding request is automatically routed to the [OpenAI Embeddings API](https://platform.openai.com/docs/guides/embeddings).
|
||||
|
||||
For example, to use OpenAI's `text-embedding-3-large` model when ingesting data, prepend the model name with `openai/` and provide your OpenAI API key in the `options` object. Any OpenAI-specific API parameters can be passed using the `options` object. This example uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/openai-upsert/" >}}
|
||||
|
||||
At query time, you can use the same model by prepending the model name with `openai/` and providing your OpenAI API key in the `options` object. This example again uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/openai-query/" >}}
|
||||
|
||||
Note that, because Qdrant does not store or cache your OpenAI API key, you need to provide it with each inference request.
|
||||
|
||||
### Cohere
|
||||
|
||||
<aside role="status">Qdrant only supports version 2 of the Cohere Embed API.</aside>
|
||||
|
||||
When you prepend a model name with `cohere/`, the embedding request is automatically routed to the [Cohere Embed API](https://docs.cohere.com/reference/embed).
|
||||
|
||||
For example, to use Cohere's multimodal `embed-v4.0` model when ingesting data, prepend the model name with `cohere/` and provide your Cohere API key in the `options` object. This example uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/cohere-upsert/" >}}
|
||||
|
||||
Note that the Cohere `embed-v4.0` model does not support passing an image as a URL. You need to provide a base64-encoded image as a Data URL.
|
||||
|
||||
At query time, you can use the same model by prepending the model name with `cohere/` and providing your Cohere API key in the `options` object. This example again uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/cohere-query/" >}}
|
||||
|
||||
Note that, because Qdrant does not store or cache your Cohere API key, you need to provide it with each inference request.
|
||||
|
||||
### Jina AI
|
||||
|
||||
When you prepend a model name with `jinaai/`, the embedding request is automatically routed to the [Jina AI Embedding API](https://jina.ai/embeddings/).
|
||||
|
||||
For example, to use Jina AI's multimodal `jina-clip-v2` model when ingesting data, prepend the model name with `jinaai/` and provide your Jina AI API key in the `options` object. This example uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/jinaai-upsert/" >}}
|
||||
|
||||
At query time, you can use the same model by prepending the model name with `jinaai/` and providing your Jina AI API key in the `options` object. This example again uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/jinaai-query/" >}}
|
||||
|
||||
Note that, because Qdrant does not store or cache your Jina AI API key, you need to provide it with each inference request
|
||||
|
||||
## Multiple Inference Operations
|
||||
|
||||
You can run multiple inference operations within a single request, even when models are hosted in different locations. This example generates three different named vectors for a single point: image embeddings using `jina-clip-v2` hosted by Jina AI, text embeddings using `all-minilm-l6-v2` hosted by Qdrant Cloud, and BM25 embeddings using the `bm25` model executed locally by the Qdrant cluster:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/multiple/" >}}
|
||||
@@ -89,6 +89,8 @@ or record-oriented equivalent:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/insert-points/list-of-points-simple/" >}}
|
||||
|
||||
### Python client optimizations
|
||||
|
||||
The Python client has additional features for loading points, which include:
|
||||
|
||||
- Parallelization
|
||||
@@ -152,6 +154,8 @@ client.upload_points(
|
||||
)
|
||||
```
|
||||
|
||||
### Idempotence
|
||||
|
||||
All APIs in Qdrant, including point loading, are idempotent.
|
||||
It means that executing the same method several times in a row is equivalent to a single execution.
|
||||
|
||||
@@ -160,6 +164,8 @@ In this case, it means that points with the same id will be overwritten when re-
|
||||
Idempotence property is useful if you use, for example, a message queue that doesn't provide an exactly-ones guarantee.
|
||||
Even with such a system, Qdrant ensures data consistency.
|
||||
|
||||
### Named vectors
|
||||
|
||||
[_Available as of v0.10.0_](#create-vector-name)
|
||||
|
||||
If the collection was created with multiple vectors, each vector data can be provided using the vector's name:
|
||||
@@ -177,6 +183,8 @@ then it is inserted with just the specified vectors. In other words, the entire
|
||||
point is replaced, and any unspecified vectors are set to null. To keep existing
|
||||
vectors unchanged and only update specified vectors, see [update vectors](#update-vectors).
|
||||
|
||||
### Sparse vectors
|
||||
|
||||
_Available as of v1.7.0_
|
||||
|
||||
Points can contain dense and sparse vectors.
|
||||
@@ -217,6 +225,16 @@ Sparse vectors must be named and can be uploaded in the same way as dense vector
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/insert-points/sparse-vectors/" >}}
|
||||
|
||||
### Inference
|
||||
|
||||
Instead of providing vectors explicitly, Qdrant can also generate vectors using a process called [inference](/documentation/inference/). Inference is the process of creating vector embeddings from text, images, or other data types using a machine learning model.
|
||||
|
||||
You can use inference in the API wherever you can use regular vectors. For example, while upserting points, you can provide the text or image and the embedding model:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/ingest/" >}}
|
||||
|
||||
Qdrant uses the model to generate the embeddings and store the point with the resulting vector.
|
||||
|
||||
## Modify points
|
||||
|
||||
To change a point, you can modify its vectors or its payload. There are several
|
||||
|
||||
@@ -166,6 +166,20 @@ To search with named vectors (available in `query` API):
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/query-points/named-vector/" >}}
|
||||
|
||||
## Inference
|
||||
|
||||
Instead of providing vectors explicitly when ingesting or querying data, Qdrant can also generate vectors using a process called [inference](/documentation/inference/). Inference is the process of creating vector embeddings from text, images, or other data types using a machine learning model.
|
||||
|
||||
You can use inference in the API wherever you can use regular vectors. For example, while upserting points, you can provide the text or image and the embedding model:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/ingest/" >}}
|
||||
|
||||
Qdrant uses the model to generate the embeddings and store the point with the resulting vector.
|
||||
|
||||
Similarly, you can use inference at query time by providing the text or image to query with and the embedding model:
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/inference/query/" >}}
|
||||
|
||||
## Datatypes
|
||||
|
||||
Newest versions of embeddings models generate vectors with very large dimentionalities.
|
||||
|
||||
+1
-1
@@ -16,7 +16,7 @@ await client.UpsertAsync(
|
||||
new() {
|
||||
Id = 1,
|
||||
Vectors = new Image() {
|
||||
Image = "https://qdrant.tech/example.png",
|
||||
Image_ = "https://qdrant.tech/example.png",
|
||||
Model = "qdrant/clip-vit-b-32-vision",
|
||||
},
|
||||
Payload = {
|
||||
|
||||
@@ -24,14 +24,14 @@ func main() {
|
||||
}
|
||||
defer client.Close()
|
||||
|
||||
_, err = client.GetPointsClient().Upsert(ctx, &qdrant.UpsertPoints{
|
||||
_, err = client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "<your-collection>",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Id: qdrant.NewIDNum(1),
|
||||
Vectors: qdrant.NewVectorsImage(&qdrant.Image{
|
||||
Image: "https://qdrant.tech/example.png",
|
||||
Model: "qdrant/clip-vit-b-32-vision",
|
||||
Image: qdrant.NewValueString("https://qdrant.tech/example.png"),
|
||||
}),
|
||||
Payload: qdrant.NewValueMap(map[string]any{
|
||||
"title": "Example image",
|
||||
|
||||
+1
-1
@@ -32,7 +32,7 @@ public class Main {
|
||||
.setVectors(
|
||||
vectors(
|
||||
Image.newBuilder()
|
||||
.setImage("https://qdrant.tech/example.png")
|
||||
.setImage(value("https://qdrant.tech/example.png"))
|
||||
.setModel("qdrant/clip-vit-b-32-vision")
|
||||
.build()))
|
||||
.putAllPayload(Map.of("title", value("Example Image")))
|
||||
|
||||
@@ -24,11 +24,11 @@ func main() {
|
||||
}
|
||||
defer client.Close()
|
||||
|
||||
_, err = client.GetPointsClient().Upsert(ctx, &qdrant.UpsertPoints{
|
||||
_, err = client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "<your-collection>",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Id: qdrant.NewIDNum(1),
|
||||
Vectors: qdrant.NewVectorsDocument(&qdrant.Document{
|
||||
Text: "Recipe for baking chocolate chip cookies",
|
||||
Model: "<the-model-to-use>",
|
||||
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet illustrates how to use the Cohere API for query-time inference on Qdrant Cloud. Instead of supplying an explicit query vector, the query provides text, along with the name of an Cohere model. When the model name is prepended with `cohere/`, the Qdrant Cloud Inference proxy uses the Cohere API to infer embeddings out of the provided text. Qdrant will search with the resulting vector. The request also shows how to pass Cohere-specific parameters to the API. In this case, the request provides the Cohere API key and the `dimensions` parameter.
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io",
|
||||
port: 6334,
|
||||
https: true,
|
||||
apiKey: "<your-api-key>"
|
||||
);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: new Document()
|
||||
{
|
||||
Model = "cohere/embed-v4.0",
|
||||
Text = "a green square",
|
||||
Options = { ["cohere-api-key"] = "<YOUR_COHERE_API_KEY>", ["output_dimension"] = 512 },
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,29 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Query(ctx, &qdrant.QueryPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Query: qdrant.NewQueryNearest(
|
||||
qdrant.NewVectorInputDocument(&qdrant.Document{
|
||||
Text: "a green square",
|
||||
Model: "cohere/embed-v4.0",
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
|
||||
"output_dimension": 512,
|
||||
}),
|
||||
}),
|
||||
),
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,13 @@
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "a green square",
|
||||
"model": "cohere/embed-v4.0",
|
||||
"options": {
|
||||
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
|
||||
"output_dimension": 512
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,34 @@
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Document;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
|
||||
client
|
||||
.queryAsync(
|
||||
Points.QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setQuery(
|
||||
nearest(
|
||||
Document.newBuilder()
|
||||
.setModel("cohere/embed-v4.0")
|
||||
.setText("a green square")
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"cohere-api-key",
|
||||
value("<YOUR_COHERE_API_KEY>"),
|
||||
"output_dimension",
|
||||
value(512)))
|
||||
.build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=models.Document(
|
||||
text="a green square",
|
||||
model="cohere/embed-v4.0",
|
||||
options={
|
||||
"cohere-api-key": "<your_cohere_api_key>",
|
||||
"output_dimension": 512
|
||||
}
|
||||
)
|
||||
)
|
||||
```
|
||||
@@ -0,0 +1,25 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Qdrant, QdrantError,
|
||||
qdrant::{Document, Query, QueryPointsBuilder, Value},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6333").build().unwrap();
|
||||
|
||||
let mut options = HashMap::<String, Value>::new();
|
||||
options.insert("cohere-api-key".to_string(), "<YOUR_COHERE_API_KEY>".into());
|
||||
options.insert("output_dimension".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(Document {
|
||||
text: "a green square".into(),
|
||||
model: "cohere/embed-v4.0".into(),
|
||||
options,
|
||||
}))
|
||||
.build(),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: {
|
||||
text: 'a green square',
|
||||
model: 'cohere/embed-v4.0',
|
||||
options: {
|
||||
'cohere-api-key': '<your_cohere_api_key>',
|
||||
output_dimension: 512,
|
||||
},
|
||||
},
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet demonstrates how to use the Cohere API for ingest-time inference on Qdrant Cloud. The example upserts a point, but instead of providing an explicit vector, the request includes text along with the name of a Cohere model. When the model name is prepended with `cohere/`, the Qdrant Cloud Inference proxy uses the Cohere API to infer embeddings out of the provided text. Qdrant will store the resulting vector. The request also shows how to pass Cohere-specific parameters to the API. In this case, the request provides the Cohere API key and the `output_dimension` parameter.
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io", port: 6334, https: true, apiKey: "<your-api-key>");
|
||||
|
||||
await client.UpsertAsync(
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Image()
|
||||
{
|
||||
Model = "cohere/embed-v4.0",
|
||||
Image_ =
|
||||
"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAoAAAAKCAYAAACNMs+9AAAAFUlEQVR42mNk+M9Qz0AEYBxVSF+FAAhKDveksOjmAAAAAElFTkSuQmCC",
|
||||
Options =
|
||||
{
|
||||
["cohere-api-key"] = "<YOUR_COHERE_API_KEY>",
|
||||
["output_dimension"] = 512,
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,32 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Vectors: qdrant.NewVectorsImage(&qdrant.Image{
|
||||
Model: "cohere/embed-v4.0",
|
||||
Image: qdrant.NewValueString("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAoAAAAKCAYAAACNMs+9AAAAFUlEQVR42mNk+M9Qz0AEYBxVSF+FAAhKDveksOjmAAAAAElFTkSuQmCC"),
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
|
||||
"output_dimension": 512,
|
||||
}),
|
||||
}),
|
||||
},
|
||||
},
|
||||
})
|
||||
```
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
```http
|
||||
PUT /collections/{collection_name}/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAoAAAAKCAYAAACNMs+9AAAAFUlEQVR42mNk+M9Qz0AEYBxVSF+FAAhKDveksOjmAAAAAElFTkSuQmCC",
|
||||
"model": "cohere/embed-v4.0",
|
||||
"options": {
|
||||
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
|
||||
"output_dimension": 512
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
```java
|
||||
import static io.qdrant.client.PointIdFactory.id;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
import static io.qdrant.client.VectorsFactory.vectors;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Image;
|
||||
import io.qdrant.client.grpc.Points.PointStruct;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
|
||||
client
|
||||
.upsertAsync(
|
||||
"{collection_name}",
|
||||
List.of(
|
||||
PointStruct.newBuilder()
|
||||
.setId(id(1))
|
||||
.setVectors(
|
||||
vectors(
|
||||
Image.newBuilder()
|
||||
.setModel("cohere/embed-v4.0")
|
||||
.setImage(
|
||||
value(
|
||||
"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAoAAAAKCAYAAACNMs+9AAAAFUlEQVR42mNk+M9Qz0AEYBxVSF+FAAhKDveksOjmAAAAAElFTkSuQmCC"))
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"cohere-api-key",
|
||||
value("<YOUR_COHERE_API_KEY>"),
|
||||
"output_dimension",
|
||||
value(512)))
|
||||
.build()))
|
||||
.build()))
|
||||
.get();
|
||||
```
|
||||
+26
@@ -0,0 +1,26 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=1,
|
||||
vector=models.Document(
|
||||
text="a green square",
|
||||
model="cohere/embed-v4.0",
|
||||
options={
|
||||
"cohere-api-key": "<your_cohere_api_key>",
|
||||
"output_dimension": 512
|
||||
}
|
||||
)
|
||||
)
|
||||
]
|
||||
)
|
||||
```
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Payload, Qdrant, QdrantError,
|
||||
qdrant::{Document, PointStruct, UpsertPointsBuilder},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build()?;
|
||||
let mut options = HashMap::new();
|
||||
options.insert("cohere-api-key".to_string(), "<YOUR_COHERE_API_KEY>".into());
|
||||
options.insert("output_dimension".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.upsert_points(UpsertPointsBuilder::new("{collection_name}",
|
||||
vec![
|
||||
PointStruct::new(1,
|
||||
Document {
|
||||
text: "Recipe for baking chocolate chip cookies requires flour, sugar, eggs, and chocolate chips.".into(),
|
||||
model: "openai/text-embedding-3-small".into(),
|
||||
options,
|
||||
},
|
||||
Payload::default())
|
||||
]).wait(true))
|
||||
.await?;
|
||||
```
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
text: 'a green square',
|
||||
model: 'cohere/embed-v4.0',
|
||||
options: {
|
||||
'cohere-api-key': '<your_cohere_api_key>',
|
||||
output_dimension: 512,
|
||||
},
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet shows how to use inference at ingest time. The example ingests a single point into a collection. Instead of providing an explicit vector, the request includes `text` and a `model`. Qdrant will use the model to infer embeddings out of the provided text and store the resulting vector.
|
||||
@@ -0,0 +1,26 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io", port: 6334, https: true, apiKey: "<your-api-key>");
|
||||
|
||||
await client.UpsertAsync(
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Dictionary<string, Vector>
|
||||
{
|
||||
["my-bm25-vector"] = new Document()
|
||||
{
|
||||
Model = "qdrant/bm25",
|
||||
Text = "Recipe for baking chocolate chip cookies",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,30 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Vectors: qdrant.NewVectorsMap(map[string]*qdrant.Vector{
|
||||
"my-bm25-vector": qdrant.NewVectorDocument(&qdrant.Document{
|
||||
Model: "qdrant/bm25",
|
||||
Text: "Recipe for baking chocolate chip cookies",
|
||||
}),
|
||||
}),
|
||||
},
|
||||
},
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,16 @@
|
||||
```http
|
||||
PUT /collections/{collection_name}/points
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"my-bm25-vector": {
|
||||
"text": "Recipe for baking chocolate chip cookies",
|
||||
"model": "qdrant/bm25"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,38 @@
|
||||
```java
|
||||
import static io.qdrant.client.PointIdFactory.id;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
import static io.qdrant.client.VectorFactory.vector;
|
||||
import static io.qdrant.client.VectorsFactory.namedVectors;
|
||||
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Image;
|
||||
import io.qdrant.client.grpc.Points.PointStruct;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
|
||||
client
|
||||
.upsertAsync(
|
||||
"{collection_name}",
|
||||
List.of(
|
||||
PointStruct.newBuilder()
|
||||
.setId(id(1))
|
||||
.setVectors(
|
||||
namedVectors(
|
||||
Map.of(
|
||||
"my-bm25-vector",
|
||||
vector(
|
||||
Document.newBuilder()
|
||||
.setModel("qdrant/bm25")
|
||||
.setText("Recipe for baking chocolate chip cookies")
|
||||
.build()))))
|
||||
.build()))
|
||||
.get();
|
||||
```
|
||||
@@ -0,0 +1,24 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=1,
|
||||
vector={
|
||||
"my-bm25-vector": models.Document(
|
||||
text="Recipe for baking chocolate chip cookies",
|
||||
model="Qdrant/bm25",
|
||||
)
|
||||
},
|
||||
)
|
||||
],
|
||||
)
|
||||
```
|
||||
@@ -0,0 +1,22 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Payload, Qdrant, QdrantError,
|
||||
qdrant::{Document, PointStruct, UpsertPointsBuilder},
|
||||
};
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build()?;
|
||||
|
||||
client
|
||||
.upsert_points(UpsertPointsBuilder::new("{collection_name}",
|
||||
vec![
|
||||
PointStruct::new(1,
|
||||
HashMap::from([("my-bm25-vector".to_string(),
|
||||
Document {
|
||||
text: "Recipe for baking chocolate chip cookies".into(),
|
||||
model: "qdrant/bm25".into(),
|
||||
..Default::default()
|
||||
}.into())]),
|
||||
Payload::default())
|
||||
]))
|
||||
.await?;
|
||||
```
|
||||
@@ -0,0 +1,19 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
'my-bm25-vector': {
|
||||
text: 'Recipe for baking chocolate chip cookies',
|
||||
model: 'Qdrant/bm25',
|
||||
},
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet illustrates how to use the Jina AI API for query-time inference on Qdrant Cloud. Instead of supplying an explicit query vector, the query provides text, along with the name of an Jina AI model. When the model name is prepended with `jinaai/`, the Qdrant Cloud Inference proxy uses the Jina AI API to infer embeddings out of the provided text. Qdrant will search with the resulting vector. The request also shows how to pass Jina AI-specific parameters to the API. In this case, the request provides the Jina AI API key and the `dimensions` parameter.
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io",
|
||||
port: 6334,
|
||||
https: true,
|
||||
apiKey: "<your-api-key>"
|
||||
);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: new Document()
|
||||
{
|
||||
Model = "jinaai/jina-clip-v2",
|
||||
Text = "Mission to Mars",
|
||||
Options = { ["jina-api-key"] = "<YOUR_JINAAI_API_KEY>", ["dimensions"] = 512 },
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,29 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Query(ctx, &qdrant.QueryPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Query: qdrant.NewQueryNearest(
|
||||
qdrant.NewVectorInputDocument(&qdrant.Document{
|
||||
Text: "Mission to Mars",
|
||||
Model: "jinaai/jina-clip-v2",
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
|
||||
"dimensions": 512,
|
||||
}),
|
||||
}),
|
||||
),
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,13 @@
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "Mission to Mars",
|
||||
"model": "jinaai/jina-clip-v2",
|
||||
"options": {
|
||||
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
|
||||
"dimensions": 512
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,33 @@
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Document;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
client
|
||||
.queryAsync(
|
||||
Points.QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setQuery(
|
||||
nearest(
|
||||
Document.newBuilder()
|
||||
.setModel("jinaai/jina-clip-v2")
|
||||
.setText("Mission to Mars")
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"jina-api-key",
|
||||
value("<YOUR_JINAAI_API_KEY>"),
|
||||
"dimensions",
|
||||
value(512)))
|
||||
.build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=models.Document(
|
||||
text="Mission to Mars",
|
||||
model="jinaai/jina-clip-v2",
|
||||
options={
|
||||
"jina-api-key": "<your_jinaai_api_key>",
|
||||
"dimensions": 512
|
||||
}
|
||||
)
|
||||
)
|
||||
```
|
||||
@@ -0,0 +1,25 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Qdrant, QdrantError,
|
||||
qdrant::{Document, Query, QueryPointsBuilder, Value},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build().unwrap();
|
||||
|
||||
let mut options = HashMap::<String, Value>::new();
|
||||
options.insert("jina-api-key".to_string(), "<YOUR_JINAAI_API_KEY>".into());
|
||||
options.insert("dimensions".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(Document {
|
||||
text: "Mission to Mars".into(),
|
||||
model: "jinaai/jina-clip-v2".into(),
|
||||
options,
|
||||
}))
|
||||
.build(),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: {
|
||||
text: 'Mission to Mars',
|
||||
model: 'jinaai/jina-clip-v2',
|
||||
options: {
|
||||
'jina-api-key': '<your_jinaai_api_key>',
|
||||
dimensions: 512,
|
||||
},
|
||||
},
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet illustrates how to use the Jina AI API for ingest-time inference on Qdrant Cloud. The example upserts a point, but instead of providing an explicit vector, the request includes text along with the name of a Jina AI model. When the model name is prepended with `jinaai/`, the Qdrant Cloud Inference proxy uses the Jina AI API to infer embeddings out of the provided text. Qdrant will store the resulting vector. The request also shows how to pass Jina AI-specific parameters to the API. In this case, the request provides the Jina AI API key and the `dimensions` parameter.
|
||||
+28
@@ -0,0 +1,28 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io",
|
||||
port: 6334,
|
||||
https: true,
|
||||
apiKey: "<your-api-key>"
|
||||
);
|
||||
|
||||
await client.UpsertAsync(
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Document()
|
||||
{
|
||||
Model = "jinaai/jina-clip-v2",
|
||||
Text = "Mission to Mars",
|
||||
Options = { ["jina-api-key"] = "<YOUR_JINAAI_API_KEY>", ["dimensions"] = 512 },
|
||||
},
|
||||
},
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,32 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Vectors: qdrant.NewVectorsImage(&qdrant.Image{
|
||||
Model: "jinaai/jina-clip-v2",
|
||||
Image: qdrant.NewValueString("https://qdrant.tech/example.png"),
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
|
||||
"dimensions": 512,
|
||||
}),
|
||||
}),
|
||||
},
|
||||
},
|
||||
})
|
||||
```
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
```http
|
||||
PUT /collections/{collection_name}/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"image": "https://qdrant.tech/example.png",
|
||||
"model": "jinaai/jina-clip-v2",
|
||||
"options": {
|
||||
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
|
||||
"dimensions": 512
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
+39
@@ -0,0 +1,39 @@
|
||||
```java
|
||||
import static io.qdrant.client.PointIdFactory.id;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
import static io.qdrant.client.VectorsFactory.vectors;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Image;
|
||||
import io.qdrant.client.grpc.Points.PointStruct;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
|
||||
client
|
||||
.upsertAsync(
|
||||
"{collection_name}",
|
||||
List.of(
|
||||
PointStruct.newBuilder()
|
||||
.setId(id(1))
|
||||
.setVectors(
|
||||
vectors(
|
||||
Image.newBuilder()
|
||||
.setModel("jinaai/jina-clip-v2")
|
||||
.setImage(value("https://qdrant.tech/example.png"))
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"jina-api-key",
|
||||
value("<YOUR_JINAAI_API_KEY>"),
|
||||
"dimensions",
|
||||
value(512)))
|
||||
.build()))
|
||||
.build()))
|
||||
.get();
|
||||
```
|
||||
+26
@@ -0,0 +1,26 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=1,
|
||||
vector=models.Image(
|
||||
image="https://qdrant.tech/example.png",
|
||||
model="jinaai/jina-clip-v2",
|
||||
options={
|
||||
"jina-api-key": "<your_jinaai_api_key>",
|
||||
"dimensions": 512
|
||||
}
|
||||
)
|
||||
)
|
||||
]
|
||||
)
|
||||
```
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Payload, Qdrant, QdrantError,
|
||||
qdrant::{Image, PointStruct, UpsertPointsBuilder},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build()?;
|
||||
let mut options = HashMap::new();
|
||||
options.insert("jina-api-key".to_string(), "<YOUR_JINAAI_API_KEY>".into());
|
||||
options.insert("dimensions".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.upsert_points(UpsertPointsBuilder::new("{collection_name}",
|
||||
vec![
|
||||
PointStruct::new(1,
|
||||
Image {
|
||||
image: Some("https://qdrant.tech/example.png".into()),
|
||||
model: "jinaai/jina-clip-v2".into(),
|
||||
options,
|
||||
},
|
||||
Payload::default())
|
||||
]).wait(true))
|
||||
.await?;
|
||||
```
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
image: 'https://qdrant.tech/example.png',
|
||||
model: 'jinaai/jina-clip-v2',
|
||||
options: {
|
||||
'jina-api-key': '<your_jinaai_api_key>',
|
||||
dimensions: 512,
|
||||
},
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet shows how to run multiple inference operations within a single request, even when models are hosted in different locations. The request generates three different named vectors for a single point: image embeddings using `jina-clip-v2` hosted by Jina AI, text embeddings using `all-minilm-l6-v2` hosted by Qdrant Cloud, and BM25 embeddings using the `bm25` model executed locally by the Qdrant cluster.
|
||||
@@ -0,0 +1,33 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io", port: 6334, https: true, apiKey: "<your-api-key>");
|
||||
|
||||
await client.UpsertAsync(
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Dictionary<string, Vector>
|
||||
{
|
||||
["image"] = new Image()
|
||||
{
|
||||
Model = "jinaai/jina-clip-v2",
|
||||
Image_ = "https://qdrant.tech/example.png",
|
||||
Options = { ["jina-api-key"] = "<YOUR_JINAAI_API_KEY>", ["dimensions"] = 512 },
|
||||
},
|
||||
["text"] = new Document()
|
||||
{
|
||||
Model = "sentence-transformers/all-minilm-l6-v2",
|
||||
Text = "Mars, the red planet",
|
||||
},
|
||||
["bm25"] = new Document() { Model = "qdrant/bm25", Text = "Mars, the red planet" },
|
||||
},
|
||||
},
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,42 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Vectors: qdrant.NewVectorsMap(map[string]*qdrant.Vector{
|
||||
"image": qdrant.NewVectorImage(&qdrant.Image{
|
||||
Model: "jinaai/jina-clip-v2",
|
||||
Image: qdrant.NewValueString("https://qdrant.tech/example.png"),
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
|
||||
"dimensions": 512,
|
||||
}),
|
||||
}),
|
||||
"text": qdrant.NewVectorDocument(&qdrant.Document{
|
||||
Model: "sentence-transformers/all-minilm-l6-v2",
|
||||
Text: "Mars, the red planet",
|
||||
}),
|
||||
"my-bm25-vector": qdrant.NewVectorDocument(&qdrant.Document{
|
||||
Model: "qdrant/bm25",
|
||||
Text: "Recipe for baking chocolate chip cookies",
|
||||
}),
|
||||
}),
|
||||
},
|
||||
},
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,28 @@
|
||||
```http
|
||||
PUT /collections/{collection_name}/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"image": {
|
||||
"image": "https://qdrant.tech/example.png",
|
||||
"model": "jinaai/jina-clip-v2",
|
||||
"options": {
|
||||
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
|
||||
"dimensions": 512
|
||||
}
|
||||
},
|
||||
"text": {
|
||||
"text": "Mars, the red planet",
|
||||
"model": "sentence-transformers/all-minilm-l6-v2"
|
||||
},
|
||||
"bm25": {
|
||||
"text": "Mars, the red planet",
|
||||
"model": "qdrant/bm25"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,56 @@
|
||||
```java
|
||||
import static io.qdrant.client.PointIdFactory.id;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
import static io.qdrant.client.VectorFactory.vector;
|
||||
import static io.qdrant.client.VectorsFactory.namedVectors;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Document;
|
||||
import io.qdrant.client.grpc.Points.Image;
|
||||
import io.qdrant.client.grpc.Points.PointStruct;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
|
||||
client
|
||||
.upsertAsync(
|
||||
"{collection_name}",
|
||||
List.of(
|
||||
PointStruct.newBuilder()
|
||||
.setId(id(1))
|
||||
.setVectors(
|
||||
namedVectors(
|
||||
Map.of(
|
||||
"image",
|
||||
vector(
|
||||
Image.newBuilder()
|
||||
.setModel("jinaai/jina-clip-v2")
|
||||
.setImage(value("https://qdrant.tech/example.png"))
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"jina-api-key",
|
||||
value("<YOUR_JINAAI_API_KEY>"),
|
||||
"dimensions",
|
||||
value(512)))
|
||||
.build()),
|
||||
"text",
|
||||
vector(
|
||||
Document.newBuilder()
|
||||
.setModel("sentence-transformers/all-minilm-l6-v2")
|
||||
.setText("Mars, the red planet")
|
||||
.build()),
|
||||
"bm25",
|
||||
vector(
|
||||
Document.newBuilder()
|
||||
.setModel("qdrant/bm25")
|
||||
.setText("Mars, the red planet")
|
||||
.build()))))
|
||||
.build()))
|
||||
.get();
|
||||
```
|
||||
@@ -0,0 +1,37 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=1,
|
||||
vector={
|
||||
"image": models.Image(
|
||||
image="https://qdrant.tech/example.png",
|
||||
model="jinaai/jina-clip-v2",
|
||||
options={
|
||||
"jina-api-key": "<your_jinaai_api_key>",
|
||||
"dimensions": 512
|
||||
},
|
||||
),
|
||||
"text": models.Document(
|
||||
text="Mars, the red planet",
|
||||
model="sentence-transformers/all-minilm-l6-v2",
|
||||
),
|
||||
"bm25": models.Document(
|
||||
text="Mars, the red planet",
|
||||
model="Qdrant/bm25",
|
||||
),
|
||||
},
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
```
|
||||
@@ -0,0 +1,51 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Payload, Qdrant, QdrantError,
|
||||
qdrant::{Document, PointStruct, UpsertPointsBuilder, Vectors},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build()?;
|
||||
|
||||
let mut jina_options = HashMap::new();
|
||||
jina_options.insert("jina-api-key".to_string(), "<YOUR_JINAAI_API_KEY>".into());
|
||||
jina_options.insert("dimensions".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.upsert_points(
|
||||
UpsertPointsBuilder::new(
|
||||
"{collection_name}",
|
||||
vec![PointStruct::new(
|
||||
1,
|
||||
NamedVectors::default()
|
||||
.add_vector(
|
||||
"image",
|
||||
Image {
|
||||
image: Some("https://qdrant.tech/example.png".into()),
|
||||
model: "jinaai/jina-clip-v2".into(),
|
||||
options: jina_options,
|
||||
},
|
||||
)
|
||||
.add_vector(
|
||||
"text",
|
||||
Document {
|
||||
text: "Mars, the red planet".into(),
|
||||
model: "sentence-transformers/all-minilm-l6-v2".into(),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.add_vector(
|
||||
"bm25",
|
||||
Document {
|
||||
text: "How to bake cookies?".into(),
|
||||
model: "qdrant/bm25".into(),
|
||||
..Default::default()
|
||||
},
|
||||
),
|
||||
Payload::default(),
|
||||
)],
|
||||
)
|
||||
.wait(true),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
image: {
|
||||
image: 'https://qdrant.tech/example.png',
|
||||
model: 'jinaai/jina-clip-v2',
|
||||
options: {
|
||||
'jina-api-key': '<your_jinaai_api_key>',
|
||||
dimensions: 512,
|
||||
},
|
||||
},
|
||||
text: {
|
||||
text: 'Mars, the red planet',
|
||||
model: 'sentence-transformers/all-minilm-l6-v2',
|
||||
},
|
||||
bm25: {
|
||||
text: 'Mars, the red planet',
|
||||
model: 'Qdrant/bm25',
|
||||
},
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet illustrates how to use the OpenAI API for query-time inference on Qdrant Cloud. Instead of supplying an explicit query vector, the query provides text, along with the name of an OpenAI model. When the model name is prepended with `openai/`, the Qdrant Cloud Inference proxy uses the OpenAI API to infer embeddings out of the provided text. Qdrant will search with the resulting vector. The request also shows how to pass OpenAI-specific parameters to the API. In this case, the request provides the OpenAI API key and the `dimensions` parameter.
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io",
|
||||
port: 6334,
|
||||
https: true,
|
||||
apiKey: "<your-api-key>"
|
||||
);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: new Document()
|
||||
{
|
||||
Model = "openai/text-embedding-3-large",
|
||||
Text = "How to bake cookies?",
|
||||
Options = { ["openai-api-key"] = "<YOUR_OPENAI_API_KEY>", ["dimensions"] = 512 },
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,29 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Query(ctx, &qdrant.QueryPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Query: qdrant.NewQueryNearest(
|
||||
qdrant.NewVectorInputDocument(&qdrant.Document{
|
||||
Model: "openai/text-embedding-3-large",
|
||||
Text: "How to bake cookies?",
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
|
||||
"dimensions": 512,
|
||||
}),
|
||||
}),
|
||||
),
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,13 @@
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "How to bake cookies?",
|
||||
"model": "openai/text-embedding-3-large",
|
||||
"options": {
|
||||
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
|
||||
"dimensions": 512
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,33 @@
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Document;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
client
|
||||
.queryAsync(
|
||||
Points.QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setQuery(
|
||||
nearest(
|
||||
Document.newBuilder()
|
||||
.setModel("openai/text-embedding-3-large")
|
||||
.setText("How to bake cookies?")
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"openai-api-key",
|
||||
value("<YOUR_OPENAI_API_KEY>"),
|
||||
"dimensions",
|
||||
value(512)))
|
||||
.build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=models.Document(
|
||||
text="How to bake cookies?",
|
||||
model="openai/text-embedding-3-large",
|
||||
options={
|
||||
"openai-api-key": "<your_openai_api_key>",
|
||||
"dimensions": 512
|
||||
}
|
||||
)
|
||||
)
|
||||
```
|
||||
@@ -0,0 +1,25 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Qdrant, QdrantError,
|
||||
qdrant::{Document, Query, QueryPointsBuilder, Value},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build().unwrap();
|
||||
|
||||
let mut options = HashMap::<String, Value>::new();
|
||||
options.insert("openai-api-key".to_string(), "<YOUR_OPENAI_API_KEY>".into());
|
||||
options.insert("dimensions".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(Document {
|
||||
text: "How to bake cookies?".into(),
|
||||
model: "openai/text-embedding-3-large".into(),
|
||||
options,
|
||||
}))
|
||||
.build(),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: {
|
||||
text: 'How to bake cookies?',
|
||||
model: 'openai/text-embedding-3-large',
|
||||
options: {
|
||||
'openai-api-key': '<your_openai_api_key>',
|
||||
dimensions: 512,
|
||||
},
|
||||
},
|
||||
});
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet illustrates how to use the OpenAI API for ingest-time inference on Qdrant Cloud. The example upserts a point, but instead of providing an explicit vector, the request includes text along with the name of an OpenAI model. When the model name is prepended with `openai/`, the Qdrant Cloud Inference proxy uses the OpenAI API to infer embeddings out of the provided text. Qdrant will store the resulting vector. The request also shows how to pass OpenAI-specific parameters to the API. In this case, the request provides the OpenAI API key and the `dimensions` parameter.
|
||||
+24
@@ -0,0 +1,24 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io", port: 6334, https: true, apiKey: "<your-api-key>");
|
||||
|
||||
await client.UpsertAsync(
|
||||
collectionName: "{collection_name}",
|
||||
points: new List<PointStruct>
|
||||
{
|
||||
new()
|
||||
{
|
||||
Id = 1,
|
||||
Vectors = new Document()
|
||||
{
|
||||
Model = "openai/text-embedding-3-large",
|
||||
Text = "Recipe for baking chocolate chip cookies",
|
||||
Options = { ["openai-api-key"] = "<YOUR_OPENAI_API_KEY>", ["dimensions"] = 512 },
|
||||
},
|
||||
},
|
||||
}
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,32 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Upsert(ctx, &qdrant.UpsertPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Points: []*qdrant.PointStruct{
|
||||
{
|
||||
Id: qdrant.NewIDNum(uint64(1)),
|
||||
Vectors: qdrant.NewVectorsDocument(&qdrant.Document{
|
||||
Model: "openai/text-embedding-3-large",
|
||||
Text: "Recipe for baking chocolate chip cookies",
|
||||
Options: qdrant.NewValueMap(map[string]any{
|
||||
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
|
||||
"dimensions": 512,
|
||||
}),
|
||||
}),
|
||||
},
|
||||
},
|
||||
})
|
||||
```
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
```http
|
||||
PUT /collections/{collection_name}/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"text": "Recipe for baking chocolate chip cookies",
|
||||
"model": "openai/text-embedding-3-large",
|
||||
"options": {
|
||||
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
|
||||
"dimensions": 512
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
+39
@@ -0,0 +1,39 @@
|
||||
```java
|
||||
import static io.qdrant.client.PointIdFactory.id;
|
||||
import static io.qdrant.client.ValueFactory.value;
|
||||
import static io.qdrant.client.VectorsFactory.vectors;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Document;
|
||||
import io.qdrant.client.grpc.Points.PointStruct;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
|
||||
client
|
||||
.upsertAsync(
|
||||
"{collection_name}",
|
||||
List.of(
|
||||
PointStruct.newBuilder()
|
||||
.setId(id(1))
|
||||
.setVectors(
|
||||
vectors(
|
||||
Document.newBuilder()
|
||||
.setModel("openai/text-embedding-3-large")
|
||||
.setText("Recipe for baking chocolate chip cookies")
|
||||
.putAllOptions(
|
||||
Map.of(
|
||||
"openai-api-key",
|
||||
value("<YOUR_OPENAI_API_KEY>"),
|
||||
"dimensions",
|
||||
value(512)))
|
||||
.build()))
|
||||
.build()))
|
||||
.get();
|
||||
```
|
||||
+26
@@ -0,0 +1,26 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=1,
|
||||
vector=models.Document(
|
||||
text="Recipe for baking chocolate chip cookies",
|
||||
model="openai/text-embedding-3-large",
|
||||
options={
|
||||
"openai-api-key": "<your_openai_api_key>",
|
||||
"dimensions": 512
|
||||
}
|
||||
)
|
||||
)
|
||||
]
|
||||
)
|
||||
```
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Payload, Qdrant, QdrantError,
|
||||
qdrant::{Document, PointStruct, UpsertPointsBuilder},
|
||||
};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build()?;
|
||||
let mut options = HashMap::new();
|
||||
options.insert("openai-api-key".to_string(), "<YOUR_OPENAI_API_KEY>".into());
|
||||
options.insert("dimensions".to_string(), 512.into());
|
||||
|
||||
client
|
||||
.upsert_points(UpsertPointsBuilder::new("{collection_name}",
|
||||
vec![
|
||||
PointStruct::new(1,
|
||||
Document {
|
||||
text: "Recipe for baking chocolate chip cookies".into(),
|
||||
model: "openai/text-embedding-3-large".into(),
|
||||
options,
|
||||
},
|
||||
Payload::default())
|
||||
]).wait(true))
|
||||
.await?;
|
||||
```
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.upsert("{collection_name}", {
|
||||
points: [
|
||||
{
|
||||
id: 1,
|
||||
vector: {
|
||||
text: 'Recipe for baking chocolate chip cookies',
|
||||
model: 'openai/text-embedding-3-large',
|
||||
options: {
|
||||
'openai-api-key': '<your_openai_api_key>',
|
||||
dimensions: 512,
|
||||
},
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
```
|
||||
@@ -0,0 +1 @@
|
||||
This code snippet shows how to use inference at query time. The example queries a collection. Instead of providing an explicit query vector, the request includes `text` and a `model`. Qdrant will use the model to infer embeddings out of the provided text and search with the resulting vector.
|
||||
@@ -0,0 +1,17 @@
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient(
|
||||
host: "xyz-example.qdrant.io",
|
||||
port: 6334,
|
||||
https: true,
|
||||
apiKey: "<your-api-key>"
|
||||
);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: new Document() { Model = "qdrant/bm25", Text = "How to bake cookies?" },
|
||||
usingVector: "my-bm25-vector"
|
||||
);
|
||||
```
|
||||
@@ -0,0 +1,26 @@
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"time"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "xyz-example.qdrant.io",
|
||||
Port: 6334,
|
||||
APIKey: "<paste-your-api-key-here>",
|
||||
UseTLS: true,
|
||||
})
|
||||
|
||||
client.Query(ctx, &qdrant.QueryPoints{
|
||||
CollectionName: "{collection_name}",
|
||||
Query: qdrant.NewQueryNearest(
|
||||
qdrant.NewVectorInputDocument(&qdrant.Document{
|
||||
Model: "qdrant/bm25",
|
||||
Text: "How to bake cookies?",
|
||||
}),
|
||||
),
|
||||
Using: qdrant.PtrOf("my-bm25-vector"),
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,10 @@
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "How to bake cookies?",
|
||||
"model": "qdrant/bm25"
|
||||
},
|
||||
"using": "my-bm25-vector"
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,27 @@
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points;
|
||||
import io.qdrant.client.grpc.Points.Document;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(
|
||||
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
|
||||
.withApiKey("<your-api-key")
|
||||
.build());
|
||||
client
|
||||
.queryAsync(
|
||||
Points.QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setQuery(
|
||||
nearest(
|
||||
Document.newBuilder()
|
||||
.setModel("qdrant/bm25")
|
||||
.setText("How to bake cookies?")
|
||||
.build()))
|
||||
.setUsing("my-bm25-vector")
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
@@ -0,0 +1,18 @@
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
url="https://xyz-example.qdrant.io:6333",
|
||||
api_key="<your-api-key>",
|
||||
cloud_inference=True
|
||||
)
|
||||
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=models.Document(
|
||||
text="How to bake cookies?",
|
||||
model="Qdrant/bm25",
|
||||
),
|
||||
using="my-bm25-vector",
|
||||
)
|
||||
```
|
||||
@@ -0,0 +1,21 @@
|
||||
```rust
|
||||
use qdrant_client::{
|
||||
Qdrant, QdrantError,
|
||||
qdrant::{Document, Query, QueryPointsBuilder},
|
||||
};
|
||||
|
||||
let client = Qdrant::from_url("<your-qdrant-url>").build().unwrap();
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(Document {
|
||||
text: "How to bake cookies?".into(),
|
||||
model: "qdrant/bm25".into(),
|
||||
..Default::default()
|
||||
}))
|
||||
.using("my-bm25-vector")
|
||||
.build(),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
@@ -0,0 +1,13 @@
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: {
|
||||
text: 'How to bake cookies?',
|
||||
model: 'qdrant/bm25',
|
||||
},
|
||||
using: 'my-bm25-vector',
|
||||
});
|
||||
```
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 195 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 1.9 MiB |
Binary file not shown.
|
After Width: | Height: | Size: 1.1 MiB |
Reference in New Issue
Block a user