Add MRL to Inference docs

This commit is contained in:
Abdon Pijpelink
2025-12-02 17:54:11 +01:00
parent 2eb4cc3d79
commit 4bf27de7b9
5 changed files with 65 additions and 1 deletions
@@ -245,4 +245,20 @@ Note that, because Qdrant does not store or cache your Jina AI API key, you need
You can run multiple inference operations within a single request, even when models are hosted in different locations. This example generates three different named vectors for a single point: image embeddings using `jina-clip-v2` hosted by Jina AI, text embeddings using `all-minilm-l6-v2` hosted by Qdrant Cloud, and BM25 embeddings using the `bm25` model executed locally by the Qdrant cluster:
{{< code-snippet path="/documentation/headless/snippets/inference/multiple/" >}}
{{< code-snippet path="/documentation/headless/snippets/inference/multiple/" >}}
When specifying multiple identical inference objects in a single request, the inference proxy executes inference only once and reuses the resulting embeddings. This optimization is particularly beneficial when working with external model providers, as it reduces both latency and cost.
## Reduce Vector Dimensionality with Matryoshka Models
[Matryoshka Representation Learning](https://arxiv.org/abs/2205.13147) (MRL) is a technique used to train embedding models to produce vectors that can be reduced in size with minimal loss of information. On Qdrant Cloud, for supported models, you can specify the `mrl` parameter in the `options` object to reduce the vector size to the desired dimension. For example:
{{< code-snippet path="/documentation/headless/snippets/inference/mrl/" >}}
By using the `mrl` option, vectors are reduced in size by the Qdrant Cloud inference proxy. This is beneficial when you are using an external model provider and need multiple vector sizes. Instead of making separate requests to the external API for each vector size, the proxy makes a single request for the original full-sized vector and then reduces it to the requested smaller size, reducing latency and cost.
A good use case for MRL is [prefetching](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries) with smaller vectors, followed by re-scoring with original-sized vectors, effectively balancing speed and accuracy. For example:
{{< code-snippet path="/documentation/headless/snippets/inference/mrl-multi-stage/" >}}
This example first prefetches 1000 candidates using a 64-dimensional reduced vector called `small` and then re-scores them using the original full-size vector called `large` to return the top 10 most relevant results.
@@ -0,0 +1 @@
This code snippet illustrates how to use smaller vectors for the initial prefetching of candidates from a large collection, followed by re-scoring with the original-sized vectors to improve accuracy, combined with inference. For the smaller vector, it employs Matryoshka Representation Learning (MRL) to reduce the dimensionality of embeddings by specifying the `mrl` parameter in the `options` object.
@@ -0,0 +1,26 @@
```http
POST /collections/{collection_name}/points/query
{
"prefetch": {
"query": {
"text": "How to bake cookies?",
"model": "openai/text-embedding-3-small",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
"mrl": 64
}
},
"using": "small",
"limit": 1000
},
"query": {
"text": "How to bake cookies?",
"model": "openai/text-embedding-3-small",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>"
}
},
"using": "large",
"limit": 10
}
```
@@ -0,0 +1 @@
This code snippet illustrates how to reduce the dimensionality of embeddings using Matryoshka Representation Learning (MRL) when using inference. It demonstrates how to insert a point into a Qdrant collection with a reduced-size vector by specifying the `mrl` parameter in the `options` object.
@@ -0,0 +1,20 @@
```http
PUT /collections/{collection_name}/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"small": {
"text": "Recipe for baking chocolate chip cookies",
"model": "openai/text-embedding-3-small",
"options": {
"openai-api-key": "<YOUR_OPENAI_API_KEY>",
"mrl": 64
}
}
}
}
]
}
```