Files
landing_page/qdrant-landing/content/documentation/inference/matryoshka-models.md
T
Abdon PijpelinkandClaude Sonnet 4.6 478b96554f Restructure inference docs (#2225)
* Break Inference page into several pages

* Make all inference code snippets testable and clean up

* Make more snippets testable

* Edits

* Document automatic query and passage prefix injection in Cloud Inference

Qdrant Cloud Inference silently applies model-specific prefixes (e.g.
"query: "/"passage: " for E5, BGE-style instruction prefix for BGE/mxbai/
Snowflake arctic-embed) so users don't need to manage them manually.
Add a section explaining this behavior, the idempotency guarantee, and
the scope (Qdrant-hosted models only; external providers handle their own).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Document short query optimization in Cloud Inference

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Update links

* Expand on external provider API key usage

* Add section about external provider API keys

* Default to header for external API keys

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 08:11:03 +02:00

20 lines
1.7 KiB
Markdown

---
title: Matryoshka Models
weight: 50
---
# Reduce Vector Dimensionality with Matryoshka Models
[Matryoshka Representation Learning](https://arxiv.org/abs/2205.13147) (MRL) is a technique used to train embedding models to produce vectors that can be reduced in size with minimal loss of information. On Qdrant Cloud, for supported models, you can specify the `mrl` parameter in the `options` object to reduce the vector size to the desired dimension.
MRL on Qdrant Cloud helps minimize costs and latency when you need multiple sizes of the same vector. Instead of making several inference requests for each vector size, the inference service only generates embeddings for the full-sized vector and then reduces the vector to each requested smaller size.
The following example demonstrates how to insert a point into a collection with both the original full-size vector (`large`) and a reduced-size vector (`small`):
{{< code-snippet path="/documentation/headless/snippets/inference/mrl/" >}}
Note that, even though the request contains two inference objects, Qdrant Cloud's inference service only makes one inference request to the OpenAI API, saving one round trip and reducing costs.
A good use case for MRL is [prefetching](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries) with smaller vectors, followed by re-scoring with the original-sized vectors, effectively balancing speed and accuracy. This example first prefetches 1000 candidates using a 64-dimensional reduced vector (`small`) and then re-scores them using the original full-size vector (`large`) to return the top 10 most relevant results:
{{< code-snippet path="/documentation/headless/snippets/inference/mrl-multi-stage/" >}}