[v1.15.0] MMR docs (#1754)

* add mmr to hybrid queries section

* update to new interface

* proofread

* update for renaming `lambda` into `diversity`

* update for `candidates_limit`

* add formula

* fix rust snippet

* docs: Simplified go.md

* docs: Simplify csharp.md

* docs: Simplified java.md

* output is regular scores

---------

Co-authored-by: Anush <anushshetty90@gmail.com>
This commit is contained in:
Luis Cossío
2025-07-18 13:24:16 +02:00
committed by GitHub
co-authored by Anush
parent ef31a4afe0
commit 33c9c18f9a
9 changed files with 181 additions and 26 deletions
@@ -8,7 +8,7 @@ hideInSidebar: false # Optional. If true, the page will not be shown in the side
# Hybrid and Multi-Stage Queries
*Available as of v1.10.0*
_Available as of v1.10.0_
With the introduction of [many named vectors per point](/documentation/concepts/vectors/#named-vectors), there are use-cases when the best search is obtained by combining multiple queries,
or by performing the search in more than one stage.
@@ -18,6 +18,7 @@ Qdrant has a flexible and universal interface to make this possible, called `Que
The main component for making the combinations of queries possible is the `prefetch` parameter, which enables making sub-requests.
Specifically, whenever a query has at least one prefetch, Qdrant will:
1. Perform the prefetch query (or queries),
2. Apply the main query over the results of its prefetch(es).
@@ -46,7 +47,7 @@ Reciprocal Rank Fusion
- `dbsf` -
<a href=https://medium.com/plain-simple-software/distribution-based-score-fusion-dbsf-a-new-approach-to-vector-search-ranking-f87c37488b18 target="_blank">
Distribution-Based Score Fusion
</a> *(available as of v1.11.0)*
</a> _(available as of v1.11.0)_
Normalizes the scores of the points in each query, using the mean +/- the 3rd standard deviation as limits, and then sums the scores of the same point across different queries.
@@ -62,14 +63,14 @@ In many cases, the usage of a larger vector representation gives more accurate s
Splitting the search into two stages is a known technique:
* First, use a smaller and cheaper representation to get a large list of candidates.
* Then, re-score the candidates using the larger and more accurate representation.
- First, use a smaller and cheaper representation to get a large list of candidates.
- Then, re-score the candidates using the larger and more accurate representation.
There are a few ways to build search architectures around this idea:
* The quantized vectors as a first stage, and the full-precision vectors as a second stage.
* Leverage Matryoshka Representation Learning (<a href=https://arxiv.org/abs/2205.13147 target="_blank">MRL</a>) to generate candidate vectors with a shorter vector, and then refine them with a longer one.
* Use regular dense vectors to pre-fetch the candidates, and then re-score them with a multi-vector model like <a href=https://arxiv.org/abs/2112.01488 target="_blank">ColBERT</a>.
- The quantized vectors as a first stage, and the full-precision vectors as a second stage.
- Leverage Matryoshka Representation Learning (<a href=https://arxiv.org/abs/2205.13147 target="_blank">MRL</a>) to generate candidate vectors with a shorter vector, and then refine them with a longer one.
- Use regular dense vectors to pre-fetch the candidates, and then re-score them with a multi-vector model like <a href=https://arxiv.org/abs/2112.01488 target="_blank">ColBERT</a>.
To get the best of all worlds, Qdrant has a convenient interface to perform the queries in stages,
such that the coarse results are fetched first, and then they are refined later with larger vectors.
@@ -88,6 +89,28 @@ It is possible to combine all the above techniques in a single query:
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring-multistage/" >}}
### Maximal Marginal Relevance (MMR)
_Available as of v1.15.0_
A useful algorithm to improve the diversity of the results is [Maximal Marginal Relevance (MMR)](https://www.cs.cmu.edu/~jgc/publication/The_Use_MMR_Diversity_Based_LTMIR_1998.pdf). It excels when the dataset has many redundant or very similar points for a query.
MMR selects candidates iteratively, starting with the most relevant point (higher similarity to the query). For each next point, it selects the one that hasn't been chosen yet which has the best combination of relevance and higher separation to the already selected points.
$$
MMR = \arg \max_{D_i \in R\setminus S}[\lambda sim(D_i, Q) - (1 - \lambda)\max_{D_j \in S}sim(D_i, D_j)]
$$
<figcaption align="center">Where $R$ is the candidates set, $S$ is the selected set, $Q$ is the query vector, $sim$ is the similarity function, and $\lambda = 1 - diversity$.</figcaption>
<br>
This is implemented in Qdrant as a parameter of a nearest neighbors query. You define the vector to get the nearest candidates, and a `diversity` parameter which controls the balance between relevance (0.0) and diversity (1.0).
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-mmr/" >}}
**Caveat:** Since MMR ranks one point at a time, the scores produced by MMR in Qdrant refer to the similarity to the query vector. This means that the response will not be ordered by score, but rather by the order of selection of MMR.
## Score boosting
_Available as of v1.14.0_
@@ -108,6 +131,7 @@ Pseudocode would be something like:
`score = score + (is_title * 0.5) + (is_content * 0.25)`
Query API can rescore points with custom formulas. They can be based on:
- Dynamic payload values
- Conditions
- Scores of prefetches
@@ -118,6 +142,7 @@ Taking the documentation example, the request would look like this:
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-tags/" >}}
There are multiple expressions available, check the [API docs for specific details](https://api.qdrant.tech/v-1-14-x/api-reference/search/query-points#request.body.query.Query%20Interface.Query.Formula%20Query.formula).
- **constant** - A floating point number. e.g. `0.5`.
- `"$score"` - Reference to the score of the point in the prefetch. This is the same as `"$score[0]"`.
- `"$score[0]"`, `"$score[1]"`, `"$score[2]"`, ... - When using multiple prefetches, you can reference specific prefetch with the index within the array of prefetches.
@@ -154,6 +179,7 @@ If there is no variable, and no defined default, a default value of `0.0` is use
</aside>
### Boost points closer to user
Another example. Combine the score with how close the result is to a user.
Considering each point has an associated geo location, we can calculate the distance between the point and the request's location.
@@ -169,7 +195,7 @@ In this case we use a **gauss_decay** function.
For all decay functions, there are these parameters available
| Parameter | Default | Description |
| --- | --- | --- |
| ---------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `x` | N/A | The value to decay |
| `target` | 0.0 | The value at which the decay will be at its peak. For distances it is usually set at 0.0, but can be set to any value. |
| `scale` | 1.0 | The value at which the decay function will be equal to `midpoint`. This is in terms of `x` units, for example, if `x` is in meters, `scale` of 5000 means 5km. Must be a non-zero positive number |
@@ -177,28 +203,25 @@ For all decay functions, there are these parameters available
The formulas for each decay function are as follows:
<iframe src="https://www.desmos.com/calculator/idv5hknwb1?embed" width="600" height="400" style="border: 1px solid #ccc" frameborder=0 class="mx-auto d-block"></iframe>
#### Decay functions
**`lin_decay`** (green), range: `[0, 1]`
$$ \text{lin_decay}(x) = \max\left(0,\ -\frac{\left(1-m_{idpoint}\right)}{s_{cale}}\cdot {abs}\left(x-t_{arget}\right)+1\right) $$
$$ \text{lin*decay}(x) = \max\left(0,\ -\frac{\left(1-m*{idpoint}\right)}{s*{cale}}\cdot {abs}\left(x-t*{arget}\right)+1\right) $$
**`exp_decay`** (red), range: `(0, 1]`
$$ \text{exp_decay}(x) = \exp\left(\frac{\ln\left(m_{idpoint}\right)}{s_{cale}}\cdot {abs}\left(x-t_{arget}\right)\right) $$
$$ \text{exp*decay}(x) = \exp\left(\frac{\ln\left(m*{idpoint}\right)}{s*{cale}}\cdot {abs}\left(x-t*{arget}\right)\right) $$
**`gauss_decay`** (purple), range: `(0, 1]`
$$ \text{gauss_decay}(x) = \exp\left(\frac{\ln\left(m_{idpoint}\right)}{s_{cale}^{2}}\cdot \left(x-t_{arget}\right)^{2}\right) $$
$$ \text{gauss*decay}(x) = \exp\left(\frac{\ln\left(m*{idpoint}\right)}{s*{cale}^{2}}\cdot \left(x-t*{arget}\right)^{2}\right) $$
## Grouping
*Available as of v1.11.0*
_Available as of v1.11.0_
It is possible to group results by a certain field. This is useful when you have multiple points for the same item, and you want to avoid redundancy of the same item in the results.
@@ -0,0 +1 @@
This code snippet demonstrates Maximal Marginal Relevance (MMR) query functionality. MMR is a technique used to balance relevance and diversity in search results by combining similarity scores with a diversity penalty. The `candidates_limit` parameter sets the number of nearest neighbors to consider for MMR. The `diversity` parameter controls this balance: values closer to 0.0 prioritize relevance, while values closer to 1.0 prioritize diversity. This approach helps avoid redundant results by penalizing documents that are too similar to already selected ones, making it particularly useful for recommendation systems and diverse result sets.
@@ -0,0 +1,19 @@
```csharp
using Qdrant.Client;
using Qdrant.Client.Grpc;
var client = new QdrantClient("localhost", 6334);
await client.QueryAsync(
collectionName: "{collection_name}",
query: (
new float[] { 0.01f, 0.45f, 0.67f },
new Mmr
{
Diversity = 0.5f, // 0.0 - relevance; 1.0 - diversity
CandidatesLimit = 100 // Number of candidates to preselect
}
),
limit: 10
);
```
@@ -0,0 +1,23 @@
```go
import (
"context"
"github.com/qdrant/go-client/qdrant"
)
client, err := qdrant.NewClient(&qdrant.Config{
Host: "localhost",
Port: 6334,
})
client.Query(context.Background(), &qdrant.QueryPoints{
CollectionName: "{collection_name}",
Query: qdrant.NewQueryMMR(
qdrant.NewVectorInput(0.01, 0.45, 0.67),
&qdrant.Mmr{
Diversity: qdrant.PtrOf(float32(0.5)), // 0.0 - relevance; 1.0 - diversity
CandidatesLimit: qdrant.PtrOf(uint32(100)), // num of candidates to preselect
}),
Limit: qdrant.PtrOf(uint64(10)),
})
```
@@ -0,0 +1,13 @@
```http
POST /collections/{collection_name}/points/query
{
"query": {
"nearest": [0.01, 0.45, 0.67, ...], // search vector
"mmr": {
"diversity": 0.5, // 0.0 - relevance; 1.0 - diversity
"candidates_limit": 100 // num of candidates to preselect
}
},
"limit": 10
}
```
@@ -0,0 +1,26 @@
```java
import static io.qdrant.client.QueryFactory.nearest;
import static io.qdrant.client.VectorInputFactory.vectorInput;
import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
import io.qdrant.client.grpc.Points.Mmr;
import io.qdrant.client.grpc.Points.QueryPoints;
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
client
.queryAsync(
QueryPoints.newBuilder()
.setCollectionName("{collection_name}")
.setQuery(
nearest(
vectorInput(0.01f, 0.45f, 0.67f), // <-- search vector
Mmr.newBuilder()
.setDiversity(0.5f) // 0.0 - relevance; 1.0 - diversity
.setCandidatesLimit(100) // num of candidates to preselect
.build()))
.setLimit(10)
.build())
.get();
```
@@ -0,0 +1,17 @@
```python
from qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
client.query_points(
collection_name="{collection_name}",
query=models.NearestQuery(
nearest=[0.01, 0.45, 0.67], # search vector
mmr=models.Mmr(
diversity=0.5, # 0.0 - relevance; 1.0 - diversity
candidates_limit=100, # num of candidates to preselect
)
),
limit=10,
)
```
@@ -0,0 +1,17 @@
```rust
use qdrant_client::Qdrant;
use qdrant_client::qdrant::{PrefetchQueryBuilder, Query, QueryPointsBuilder};
let client = Qdrant::from_url("http://localhost:6334").build()?;
client.query(
QueryPointsBuilder::new("{collection_name}")
.query(Query::new_nearest_with_mmr(
vec![0.01, 0.45, 0.67], // search vector
MmrBuilder::new()
.diversity(0.5) // 0.0 - relevance; 1.0 - diversity
.candidates_limit(100) // num of candidates to preselect
))
.limit(10)
).await?;
```
@@ -0,0 +1,16 @@
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({ host: "localhost", port: 6333 });
client.query("{collection_name}", {
query: {
nearest: [0.01, 0.45, 0.67, ...], // search vector
mmr: {
diversity: 0.5, // 0.0 - relevance; 1.0 - diversity
candidates_limit: 100 // num of candidates to preselect
}
},
limit: 10,
});
```