Separate external provider code snippets

This commit is contained in:
Abdon Pijpelink
2025-11-13 10:51:34 +01:00
parent 2541942c09
commit da0344e457
19 changed files with 130 additions and 146 deletions
@@ -183,7 +183,9 @@ Qdrant Cloud can act as a proxy for the APIs of three external embedding model p
This enables you to access any of the embedding models provided by these providers through the Qdrant API. This enables you to access any of the embedding models provided by these providers through the Qdrant API.
To use an external provider's embedding model, you need an API key from that provider. For example, to access OpenAI models, you need an OpenAI API key. Billing is managed directly through the external provider, based on API key usage. Refer to each external embedding model provider's website for pricing details. To use an external provider's embedding model, you need an API key from that provider. For example, to access OpenAI models, you need an OpenAI API key. Qdrant does not store or cache your API keys; they must be provided with each inference request.
When using an external embedding model, ensure that your collection has been configured for vectors with the correct dimensionality. Refer to the model's documentation for details on the output dimensions.
<aside role="status"> <aside role="status">
When using a model from an external provider, refer to the model's documentation for: When using a model from an external provider, refer to the model's documentation for:
@@ -195,31 +197,49 @@ When using a model from an external provider, refer to the model's documentation
### OpenAI ### OpenAI
When you prepend a model name with `openai/`, the embedding request is automatically routed to the [OpenAI Embeddings API](https://platform.openai.com/docs/guides/embeddings). You need to provide your OpenAI API key with each request. When you prepend a model name with `openai/`, the embedding request is automatically routed to the [OpenAI Embeddings API](https://platform.openai.com/docs/guides/embeddings).
For example, to use OpenAI's `text-embedding-3-large` model, when ingesting and querying data, prepend the model name with `openai/` and provide your OpenAI API key in the `options` object. Any OpenAI-specific API parameters can be passed using the `options` object. This example uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512: For example, to use OpenAI's `text-embedding-3-large` model when ingesting data, prepend the model name with `openai/` and provide your OpenAI API key in the `options` object. Any OpenAI-specific API parameters can be passed using the `options` object. This example uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/openai/" >}} {{< code-snippet path="/documentation/headless/snippets/inference/openai-upsert/" >}}
At query time, you can use the same model by prepending the model name with `openai/` and providing your OpenAI API key in the `options` object. This example again uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/openai-query/" >}}
Note that, because Qdrant does not store or cache your OpenAI API key, you need to provide it with each inference request.
### Cohere ### Cohere
<aside role="status">Qdrant only supports version 2 of the Cohere Embed API.</aside> <aside role="status">Qdrant only supports version 2 of the Cohere Embed API.</aside>
When you prepend a model name with `cohere/`, the embedding request is automatically routed to the [Cohere Embed API](https://docs.cohere.com/reference/embed). You need to provide your Cohere API key with each request. When you prepend a model name with `cohere/`, the embedding request is automatically routed to the [Cohere Embed API](https://docs.cohere.com/reference/embed).
For example, to use Cohere's multimodal `embed-v4.0` model, when ingesting and querying data, prepend the model name with `cohere/` and provide your Cohere API key in the `options` object. This example uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512: For example, to use Cohere's multimodal `embed-v4.0` model when ingesting data, prepend the model name with `cohere/` and provide your Cohere API key in the `options` object. This example uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/cohere/" >}} {{< code-snippet path="/documentation/headless/snippets/inference/cohere-upsert/" >}}
Note that the Cohere `embed-v4.0` model does not allow an image to be passed as a URL. You need to provide a base64-encoded image as a Data URL. Note that the Cohere `embed-v4.0` model does not support passing an image as a URL. You need to provide a base64-encoded image as a Data URL.
At query time, you can use the same model by prepending the model name with `cohere/` and providing your Cohere API key in the `options` object. This example again uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/cohere-query/" >}}
Note that, because Qdrant does not store or cache your Cohere API key, you need to provide it with each inference request.
### Jina AI ### Jina AI
When you prepend a model name with `jinaai/`, the embedding request is automatically routed to the [Jina AI Embedding API](https://jina.ai/embeddings/). You need to provide your Jina AI API key with each request. When you prepend a model name with `jinaai/`, the embedding request is automatically routed to the [Jina AI Embedding API](https://jina.ai/embeddings/).
For example, to use Jina AI's multimodal `jina-clip-v2` model, when ingesting and querying data, prepend the model name with `jinaai/` and provide your Jina AI API key in the `options` object. This example uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512: For example, to use Jina AI's multimodal `jina-clip-v2` model when ingesting data, prepend the model name with `jinaai/` and provide your Jina AI API key in the `options` object. This example uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/jinaai/" >}} {{< code-snippet path="/documentation/headless/snippets/inference/jinaai-upsert/" >}}
At query time, you can use the same model by prepending the model name with `jinaai/` and providing your Jina AI API key in the `options` object. This example again uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/jinaai-query/" >}}
Note that, because Qdrant does not store or cache your Jina AI API key, you need to provide it with each inference request
## Multiple Inference Operations ## Multiple Inference Operations
@@ -0,0 +1 @@
This code snippet illustrates how to use the Cohere API for query-time inference on Qdrant Cloud. Instead of supplying an explicit query vector, the query provides text, along with the name of an Cohere model. When the model name is prepended with `cohere/`, the Qdrant Cloud Inference proxy uses the Cohere API to infer embeddings out of the provided text. Qdrant will search with the resulting vector. The request also shows how to pass Cohere-specific parameters to the API. In this case, the request provides the Cohere API key and the `dimensions` parameter.
@@ -0,0 +1,13 @@
```http
POST /collections/<your-collection_name>/points/query
{
"query": {
"text": "a green square",
"model": "cohere/embed-v4.0",
"options": {
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
"output_dimension": 512
}
}
}
```
@@ -0,0 +1 @@
This code snippet demonstrates how to use the Cohere API for ingest-time inference on Qdrant Cloud. The example upserts a point, but instead of providing an explicit vector, the request includes text along with the name of a Cohere model. When the model name is prepended with `cohere/`, the Qdrant Cloud Inference proxy uses the Cohere API to infer embeddings out of the provided text. Qdrant will store the resulting vector. The request also shows how to pass Cohere-specific parameters to the API. In this case, the request provides the Cohere API key and the `output_dimension` parameter.
@@ -0,0 +1,18 @@
```http
PUT /collections/<your-collection_name>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAoAAAAKCAYAAACNMs+9AAAAFUlEQVR42mNk+M9Qz0AEYBxVSF+FAAhKDveksOjmAAAAAElFTkSuQmCC",
"model": "cohere/embed-v4.0",
"options": {
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
"output_dimension": 512
}
}
}
]
}
```
@@ -1 +0,0 @@
This code snippet demonstrates how to use the Cohere API for inference on Qdrant Cloud. The example first creates a collection that supports vectors with 512 dimensions. Next, a point is inserted, but instead of providing an explicit vector, the example request includes text along with the name of an Cohere model. When the model name is prepended with `cohere/`, the Qdrant Cloud Inference proxy will use the Cohere API to infer embeddings out of the provided text and store the resulting vector. The request also shows how to pass Cohere-specific parameters to the API. In this case, the request passes the Cohere API key and the `dimensions` parameter. Finally, the example shows how to use the Cohere API for query-time inference. Instead of supplying an explicit query vector, the query includes text and name of an Cohere model, as well as an Cohere API key and the Cohere-specific `dimensions` parameter. When the model name is prepended with `cohere/`, the Qdrant Cloud Inference proxy will use the Cohere API to infer embeddings out of the provided text and search with the resulting vector.
@@ -1,43 +0,0 @@
```http
// Create the collection for vectors with 512 dimensions.
PUT /collections/<your-collection_name>
{
"vectors": {
"size": 512,
"distance": "Cosine"
}
}
// Ingest a point. Provide the model name, prepended with "cohere/".
// Provide the Cohere API key in the "options" object.
PUT /collections/<your-collection_name>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAoAAAAKCAYAAACNMs+9AAAAFUlEQVR42mNk+M9Qz0AEYBxVSF+FAAhKDveksOjmAAAAAElFTkSuQmCC",
"model": "cohere/embed-v4.0",
"options": {
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
"output_dimension": 512
}
}
}
]
}
// Query the data by providing the model name, prepended with "cohere/"
// and the Cohere API key.
POST /collections/<your-collection_name>/points/query
{
"query": {
"text": "a green square",
"model": "cohere/embed-v4.0",
"options": {
"cohere-api-key": "<YOUR_COHERE_API_KEY>",
"output_dimension": 512
}
}
}
```
@@ -0,0 +1 @@
This code snippet illustrates how to use the Jina AI API for query-time inference on Qdrant Cloud. Instead of supplying an explicit query vector, the query provides text, along with the name of an Jina AI model. When the model name is prepended with `jinaai/`, the Qdrant Cloud Inference proxy uses the Jina AI API to infer embeddings out of the provided text. Qdrant will search with the resulting vector. The request also shows how to pass Jina AI-specific parameters to the API. In this case, the request provides the Jina AI API key and the `dimensions` parameter.
@@ -0,0 +1,13 @@
```http
POST /collections/<your-collection_name>/points/query
{
"query": {
"text": "Mission to Mars",
"model": "jinaai/jina-clip-v2",
"options": {
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
"dimensions": 512
}
}
}
```
@@ -0,0 +1 @@
This code snippet illustrates how to use the Jina AI API for ingest-time inference on Qdrant Cloud. The example upserts a point, but instead of providing an explicit vector, the request includes text along with the name of a Jina AI model. When the model name is prepended with `jinaai/`, the Qdrant Cloud Inference proxy uses the Jina AI API to infer embeddings out of the provided text. Qdrant will store the resulting vector. The request also shows how to pass Jina AI-specific parameters to the API. In this case, the request provides the Jina AI API key and the `dimensions` parameter.
@@ -0,0 +1,18 @@
```http
PUT /collections/<your-collection_name>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"image": "https://qdrant.tech/example.png",
"model": "jinaai/jina-clip-v2",
"options": {
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
"dimensions": 512
}
}
}
]
}
```
@@ -1 +0,0 @@
This code snippet demonstrates how to use the Jina AI API for inference on Qdrant Cloud. The example first creates a collection that supports vectors with 512 dimensions. Next, a point is inserted, but instead of providing an explicit vector, the example request includes text along with the name of an Jina AI model. When the model name is prepended with `jinaai/`, the Qdrant Cloud Inference proxy will use the Jina AI API to infer embeddings out of the provided text and store the resulting vector. The request also shows how to pass Jina AI-specific parameters to the API. In this case, the request passes the Jina AI API key and the `dimensions` parameter. Finally, the example shows how to use the Jina AI API for query-time inference. Instead of supplying an explicit query vector, the query includes text and name of an Jina AI model, as well as an Jina AI API key and the Jina AI-specific `dimensions` parameter. When the model name is prepended with `jinaai/`, the Qdrant Cloud Inference proxy will use the Jina AI API to infer embeddings out of the provided text and search with the resulting vector.
@@ -1,43 +0,0 @@
```http
// Create the collection for vectors with 512 dimensions.
PUT /collections/<your-collection_name>
{
"vectors": {
"size": 512,
"distance": "Cosine"
}
}
// Ingest a point. Provide the model name, prepended with "jinaai/".
// Provide the Jina AI API key in the "options" object.
PUT /collections/<your-collection_name>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"image": "https://qdrant.tech/example.png",
"model": "jinaai/jina-clip-v2",
"options": {
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
"dimensions": 512
}
}
}
]
}
// Query the data by providing the model name, prepended with "jinaai/"
// and the Jina AI API key.
POST /collections/<your-collection_name>/points/query
{
"query": {
"text": "Mission to Mars",
"model": "jinaai/jina-clip-v2",
"options": {
"jina-api-key": "<YOUR_JINAAI_API_KEY>",
"dimensions": 512
}
}
}
```
@@ -0,0 +1 @@
This code snippet illustrates how to use the OpenAI API for query-time inference on Qdrant Cloud. Instead of supplying an explicit query vector, the query provides text, along with the name of an OpenAI model. When the model name is prepended with `openai/`, the Qdrant Cloud Inference proxy uses the OpenAI API to infer embeddings out of the provided text. Qdrant will search with the resulting vector. The request also shows how to pass OpenAI-specific parameters to the API. In this case, the request provides the OpenAI API key and the `dimensions` parameter.
@@ -0,0 +1,13 @@
```http
POST /collections/<your-collection_name>/points/query
{
"query": {
"text": "How to bake cookies?",
"model": "openai/text-embedding-3-large",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
"dimensions": 512
}
}
}
```
@@ -0,0 +1 @@
This code snippet illustrates how to use the OpenAI API for ingest-time inference on Qdrant Cloud. The example upserts a point, but instead of providing an explicit vector, the request includes text along with the name of an OpenAI model. When the model name is prepended with `openai/`, the Qdrant Cloud Inference proxy uses the OpenAI API to infer embeddings out of the provided text. Qdrant will store the resulting vector. The request also shows how to pass OpenAI-specific parameters to the API. In this case, the request provides the OpenAI API key and the `dimensions` parameter.
@@ -0,0 +1,18 @@
```http
PUT /collections/<your-collection_name>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"text": "Recipe for baking chocolate chip cookies",
"model": "openai/text-embedding-3-large",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
"dimensions": 512
}
}
}
]
}
```
@@ -1 +0,0 @@
This code snippet demonstrates how to use the OpenAI API for inference on Qdrant Cloud. The example first creates a collection that supports vectors with 512 dimensions. Next, a point is inserted, but instead of providing an explicit vector, the example request includes text along with the name of an OpenAI model. When the model name is prepended with `openai/`, the Qdrant Cloud Inference proxy will use the OpenAI API to infer embeddings out of the provided text and store the resulting vector. The request also shows how to pass OpenAI-specific parameters to the API. In this case, the request passes the OpenAI API key and the `dimensions` parameter. Finally, the example shows how to use the OpenAI API for query-time inference. Instead of supplying an explicit query vector, the query includes text and name of an OpenAI model, as well as an OpenAI API key and the OpenAI-specific `dimensions` parameter. When the model name is prepended with `openai/`, the Qdrant Cloud Inference proxy will use the OpenAI API to infer embeddings out of the provided text and search with the resulting vector.
@@ -1,46 +0,0 @@
```http
// Configure the collection for vectors with 512 dimensions.
PUT /collections/<your-collection_name>
{
"vectors": {
"size": 512,
"distance": "Cosine"
}
}
// Ingest a point. Provide the model name, prepended with "openai/".
// Provide the OpenAI API key in the "options" object.
PUT /collections/<your-collection_name>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"text": "Recipe for baking chocolate chip cookies",
"model": "openai/text-embedding-3-large",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
"dimensions": 512
}
}
}
]
}
// Retrieve the point to see the generated embeddings
GET /collections/<your-collection_name>/points/1
// Query the data by providing the model name, prepended with "openai/"
// and the OpenAI API key.
POST /collections/<your-collection_name>/points/query
{
"query": {
"text": "How to bake cookies?",
"model": "openai/text-embedding-3-large",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
"dimensions": 512
}
}
}
```