diff --git a/qdrant-landing/content/documentation/concepts/inference.md b/qdrant-landing/content/documentation/concepts/inference.md index 0796407cc..7647fab38 100644 --- a/qdrant-landing/content/documentation/concepts/inference.md +++ b/qdrant-landing/content/documentation/concepts/inference.md @@ -251,14 +251,14 @@ When specifying multiple identical inference objects in a single request, the in ## Reduce Vector Dimensionality with Matryoshka Models -[Matryoshka Representation Learning](https://arxiv.org/abs/2205.13147) (MRL) is a technique used to train embedding models to produce vectors that can be reduced in size with minimal loss of information. On Qdrant Cloud, for supported models, you can specify the `mrl` parameter in the `options` object to reduce the vector size to the desired dimension. For example: +[Matryoshka Representation Learning](https://arxiv.org/abs/2205.13147) (MRL) is a technique used to train embedding models to produce vectors that can be reduced in size with minimal loss of information. On Qdrant Cloud, for supported models, you can specify the `mrl` parameter in the `options` object to reduce the vector size to the desired dimension. + +By using the `mrl` option, vectors are reduced in size by the Qdrant Cloud inference proxy. This is beneficial when you are using an external model provider and need multiple vector sizes. External model APIs also offer options to reduce vector size, but to request multiple sizes, you would need to make several requests to the external API, incurring cost and latency. However, with the `mrl` option, the proxy only makes a single external API request for the original full-sized vector and then reduces it to the requested smaller size, minimizing cost and latency. + +The following example demonstrates how to insert a point into a collection with both the original full-size vector (`large`) and a reduced-size vector (`small`). Even though the request contains two inference objects, Qdrant Cloud's inference proxy only makes one request to the OpenAI API: {{< code-snippet path="/documentation/headless/snippets/inference/mrl/" >}} -By using the `mrl` option, vectors are reduced in size by the Qdrant Cloud inference proxy. This is beneficial when you are using an external model provider and need multiple vector sizes. Instead of making separate requests to the external API for each vector size, the proxy makes a single request for the original full-sized vector and then reduces it to the requested smaller size, reducing latency and cost. +A good use case for MRL is [prefetching](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries) with smaller vectors, followed by re-scoring with original-sized vectors, effectively balancing speed and accuracy. This example first prefetches 1000 candidates using a 64-dimensional reduced vector (`small`) and then re-scores them using the original full-size vector (`large`) to return the top 10 most relevant results: -A good use case for MRL is [prefetching](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries) with smaller vectors, followed by re-scoring with original-sized vectors, effectively balancing speed and accuracy. For example: - -{{< code-snippet path="/documentation/headless/snippets/inference/mrl-multi-stage/" >}} - -This example first prefetches 1000 candidates using a 64-dimensional reduced vector called `small` and then re-scores them using the original full-size vector called `large` to return the top 10 most relevant results. \ No newline at end of file +{{< code-snippet path="/documentation/headless/snippets/inference/mrl-multi-stage/" >}} \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/inference/mrl/http.md b/qdrant-landing/content/documentation/headless/snippets/inference/mrl/http.md index 29f6ac33c..79fa4ce65 100644 --- a/qdrant-landing/content/documentation/headless/snippets/inference/mrl/http.md +++ b/qdrant-landing/content/documentation/headless/snippets/inference/mrl/http.md @@ -5,6 +5,13 @@ PUT /collections/{collection_name}/points?wait=true { "id": 1, "vector": { + "large": { + "text": "Recipe for baking chocolate chip cookies", + "model": "openai/text-embedding-3-small", + "options": { + "openai-api-key": "" + } + }, "small": { "text": "Recipe for baking chocolate chip cookies", "model": "openai/text-embedding-3-small",