fixes according to new visualizations

This commit is contained in:
Evgeniya Sukhodolskaya
2026-02-18 11:04:18 +01:00
parent 04318fdea5
commit ff9aa93ac1
14 changed files with 164 additions and 107 deletions
@@ -6,7 +6,7 @@ social_preview_image: /articles_data/relevance-feedback/preview/social_preview.j
preview_dir: /articles_data/relevance-feedback/preview
weight: -170
author: Evgeniya Sukhodolskaya
date: 2025-10-31T00:00:00+03:00
date: 2026-02-10T00:00:00+03:00
draft: false
keywords:
- search relevance
@@ -16,13 +16,13 @@ keywords:
category: machine-learning
---
This March, we dropped a statement-bomb in the “[Relevance Feedback in Information Retrieval](https://qdrant.tech/articles/search-feedback-loop/)” article and then went silent.
A year ago, we dropped a statement-bomb in the “[Relevance Feedback in Information Retrieval](https://qdrant.tech/articles/search-feedback-loop/)” article and then went silent.
We claimed that even though the information retrieval research field has proposed many useful mechanisms for increasing the relevance of search results, none of them made it to the neural search industry, simply because these approaches are not scalable.
Certainly, there are methods widely used to improve the relevance of retrieved results: query rewriting, for example. Yet none of the vector search solutions out there have tried to use the possibilities that come with full access to the vector search index: traversing it in the direction of relevance, instead of guessing where to shoot the query to the vector space.
To break this vector-search-native tools silence, we’re introducing FeedbackQuery <TBD link to release> — a universal, cheap and scalable method for improving the relevance recall of your search results.
To break this vector-search-native tools silence, we’re introducing [Relevance Feedback Query](https://qdrant.tech/blog/qdrant-1.17.x/#relevance-feedback-query), a universal, cheap and scalable method for improving the relevance recall of your search results.
## Vector Search Optimization
@@ -34,7 +34,7 @@ For those instruments to work (and to stick in industry), they have to be **univ
At scale, neural search is usually limited to fairly simple dense encoders (embeddings up to a few thousand dimensions), as running large models to sift through billions of data points is, well, unreasonable.
So most search quality-boosting tools aim to adjust/align/guide these simple retrievers. The guidance can come from the search user or a domain expert: rescoring via business-logic-based rules, reranking with a learned-to-rank model, and using **relevance feedback** mechanisms.
So most search quality-boosting tools aim to adjust/align/guide these simple retrievers. The guidance can come from the search end-user or a domain expert: rescoring via business-logic-based rules, reranking with a learned-to-rank model, and using **relevance feedback** mechanisms.
### What is Relevance Feedback
@@ -50,11 +50,11 @@ Based on this feedback, one of the retrieval components is adjusted: the query o
Relevance-feedback-based methods are a standard in full-text search, some of them proposed more than 50 years ago.
Yet it feels like, when it comes to modern vector search, users have to reinvent the relevance feedback wheel due to the lack of universal interfaces: prompt an agent to rewrite queries in endless, costly loops, do heuristics-based vector math and fine-tune models per request on the client side... Simply put, they have to struggle.
Yet it feels like, when it comes to modern vector search, users have to reinvent the relevance feedback wheel due to the lack of universal interfaces: prompt a search agent to rewrite queries in endless, costly loops, do heuristics-based vector math and fine-tune models per request on the client side... Simply put, they have to struggle.
### Tools of a Vector Search Engine
The goal is to fill this gap and let users get a proper relevance feedback interface to use in our vector search engine.
The goal is to fill this gap and let search end-users (doesn't matter if humans or agents) get a proper relevance feedback interface to use in our vector search engine.
So what makes a method a production-ready for vector search?
@@ -78,7 +78,7 @@ Models which require many labels or domain expert input are hard to adopt. The f
#### ...Should be Universal
**For every type of data**
Vector search is attractive because it is agnostic to the data type: images, video, audio… Not only text. Vector search native tools should be the same, operating on vectors without minding their nature. That is why query rewriting as reformulating text won’t suffice.
Vector search is attractive because it is agnostic to the data type: images, video, audio, molecules… Not only text. Vector search native tools should be the same, operating on vectors without minding their nature. That is why query rewriting as reformulating text won’t suffice.
**For every type of signal**
Some relevance feedback methods depend on clean feedback signals: strictly relevant or irrelevant documents.
@@ -116,7 +116,7 @@ And the hiker’s new “where-should-I-go” decisions based on the acquired in
So the idea is to use feedback to define a more relevant direction (**closer** to this, **further** from that) within the vector space spanned by the dataset.
That implies warping the notion of "closer" and "further" – the distance (or similarity) **scoring metric** used during the retrieval.
That implies warping the notion of "closer" and "further", the distance (or similarity) **scoring metric** used during the retrieval.
Tweaking the relevance scoring formula based on the feedback model’s signal is nothing revolutionary. What changes is that this feedback-based scoring is used during the traversal of the **entire** dataset used by the vector search engine, not just a subset of results.
@@ -126,17 +126,20 @@ So, stepping away from analogies (into an even deeper forest), the algorithm sho
1. Initial retrieval.
2. Getting a small amount of feedback, within our **context limit**. A good choice, for example, is granular judgments, pointwise scores between the query and the top retrieved documents. They are more helpful when results are not strictly relevant or irrelevant.
2. Getting a small amount of feedback (within the **context limit**) on the results of this initial retrieval.
> A **context limit** is how many top documents from the initial retrieval the feedback model scores. We want to keep it as small as possible to **save time and resources**.
A good choice for feedback format is granular judgments, aka pointwise relevance scores, between the query and the retrieved documents. They are more helpful when the retrieved documents are not strictly relevant or irrelevant.
3. Extracting relevance signals from the feedback and propagating them into a new similarity (distance) scoring formula.
3. Extracting relevance signals from this feedback and propagating them into a new similarity (distance) scoring formula.
4. Using this new feedback-based formula in the next retrieval iteration on the whole dataset of documents to traverse the vector index in the **direction of more relevance**.
4. Using this new feedback-based similarity scoring formula in the next retrieval iteration **on the whole dataset of documents** to traverse the vector index in the direction of more relevance.
Simply put, we're no longer using cosine to score similarities, but rather a formula adjusted for feedback.
## Dissecting Feedback
Let’s think about how a very small amount of feedback can be used **most effectively**, and what signals we can extract from it.
Let's talk about how a very small amount of feedback can be used most effectively, what relevance signals we can extract from it, and how exactly.
### Context Pairs
@@ -146,160 +149,189 @@ With two documents, this model could already judge which one is closer to what t
> A **context pair (positive, negative)** is two documents from the top context limit results of the initial retrieval. The positive received a higher relevance-to-query score from the feedback model, and the negative received a lower score.
{{< figure src="/articles_data/relevance-feedback/context_pairs.png" alt="The image illustrates how context pairs are formed. On the left, a ‘retriever’ column shows stacked colored blocks representing retrieved documents, with each color indicating the feedback assigned to that document. A dotted ‘context limit’ line marks which retrieved items are used for feedback. On the right, a ‘context pairs’ column shows pairs constructed from these feedback-colored items, each pair combining one more-positive item and one more-negative item, as indicated by their colors." caption="How context pairs are formed">}}
{{< figure src="/articles_data/relevance-feedback/context_pair.png" alt="Diagram titled Re-score with Feedback Model showing five documents labeled Doc 1 to Doc 5, each with two horizontal bars for Retriever Score in blue and Feedback Score in green. Doc 2 has the highest feedback score and Doc 5 has the lowest feedback score. Doc 2 and Doc 4 show higher feedback scores than retriever scores, while Doc 3 and Doc 5 show lower feedback scores. On the right, a Context Pair section shows Doc 2 labeled Positive and Doc 5 labeled Negative, with curved arrows linking these examples back to the document list" caption="How context pairs are formed">}}
### Feedback’s Confidence
### Context Pair's Confidence
Two documents nearly indistinguishable from a perspective of a feedback model give far less information than a context pair with one clearly more relevant document.
When documents in the context pair seem different to the feedback model, it is more **confident** to guide the retriever.
{{< figure src="/articles_data/relevance-feedback/confidence.png" alt="The image shows several context pairs, each represented by two horizontal colored blocks stacked together. Next to each pair is a dashed double-headed arrow labeled ‘confidence.’ The color of each arrow indicates the confidence level, with stronger color showing higher confidence and lighter color showing lower confidence. Confidence reflects how different the two documents in each context pair are, as indicated by the contrast between their colors." caption="Confidence reflects how different the documents in each context pair are." width="50%">}}
{{< figure src="/articles_data/relevance-feedback/confidence_of_context_pair.png" alt="Diagram titled Re-score with Feedback Model showing five documents labeled Doc 1 to Doc 5, each with two horizontal bars for Retriever Score in blue and Feedback Score in green. Doc 2 has the highest feedback score and Doc 5 has the lowest feedback score. Doc 2 and Doc 4 show higher feedback scores than retriever scores, while Doc 3 and Doc 5 show lower feedback scores. On the right, a Context Pair section shows Doc 2 labeled Positive and Doc 5 labeled Negative, with curved arrows linking these examples back to the document list" caption="Context pair's confidence">}}
### Direction's Delta
Then, if the retriever favors, by some **delta,** a document closer to the negative example, it has chosen the wrong direction in the vector space and should change it.
Then, if the retriever favors a document closer to the negative example by some **delta**, the retriever should most probably adjust its search direction in the vector space.
{{< figure src="/articles_data/relevance-feedback/delta.png" alt="The image shows a candidate document represented as a dark dot in the center, with dashed arrows indicating its distances to a circled positive example above and a circled negative example below. A ‘|delta|’ segment on the arrow to the positive example highlights how much more the retriever currently favors the negative example. The diagram illustrates that the retriever should adjust its direction toward candidates closer to the positive example, as shown by an additional dashed arrow." caption="The retriever currently favors the negative element of a context pair by a delta, and should select a candidate closer to the positive one">}}
{{< figure src="/articles_data/relevance-feedback/delta_as_distance.png" alt="The image shows a candidate document represented as a dark dot in the center, with dashed arrows indicating its distances to a circled positive example above and a circled negative example below. A ‘-delta’ segment on the arrow to the positive example highlights how much more the retriever currently favors the negative example. The diagram illustrates that the retriever should adjust its direction toward candidates closer to the positive example, as shown by an additional dashed arrow." caption="The point in vector space is closer (more similar) to the negative element of a context pair by a delta." width="80%">}}
## Feedback-based Scoring
With the context pair(s) at our expense, we can try the following feedback-based scoring:
With the context pair(s) at our expense, we can try the following feedback-based scoring during retrieval:
1. Let the retriever still have a say in what is relevant to the query, count in its **score**.
1. Let the retriever still have a say in what is relevant to the query, count in its **score** (to query).
2. Yet reward candidates that are closer to the positive element of the context pair based on **delta**.
3. Especially when the feedback model had **confidence** in this pair.
3. Especially when the feedback model had high **confidence** in this pair.
{{< figure src="/articles_data/relevance-feedback/relevance_feedback_scoring.png" alt="The image shows a query point and a candidate document in a vector space. A blue dashed arrow labeled ‘score’ connects the query to the candidate, representing similarity (or distance). The candidate also has dashed arrows pointing toward a circled positive example and a circled negative example, showing its distances to each. A marked ‘|delta|’ segment on the positive-direction arrow highlights how much more the retriever currently favors the negative direction. The diagram illustrates feedback-based scoring that combines the candidate’s similarity to the query with its relative distances to the positive and negative items in the context pair." caption="Feedback-based scoring based on the candidate’s distance (similarity) to the query and the context pair.">}}
{{< figure src="/articles_data/relevance-feedback/scoring.png" alt="Diagram titled Scoring with context pairs showing how candidates from a collection are scored using three components: similarity to the query, similarity to a positive example Doc 2, and similarity to a negative example Doc 5. For each candidate, horizontal bars show score to query in blue, score to positive in green, and score to negative in red. A dashed delta indicates the difference between the positive and negative scores, which rewards candidates closer to the positive example and farther from the negative one. Larger deltas represent stronger separation in relevance, especially when the context pair is known with high confidence." caption="Feedback-based scoring using candidate’s similarity to the query and the context pair.">}}
We need to combine signals in a reasonable, simple fashion, using as few parameters as we can.
> We need trainable parameters in the feedback-based scoring formula because relevance score distributions likely vary by dataset, retriever, and feedback model.
Usually what helps is looking at edge cases.
So, we need to combine signals (**score, delta, confidence**) in a reasonable, simple scoring formula with a few parameters.
Usually what helps in coming up with a reasonable formula is looking at edge cases.
### The Math of Edge Cases
**Direction’s delta is zero**
If the retriever sees the positive and the negative documents (from the context pair) as equally similar to the candidate document, there’s no direction to choose.
If the retriever sees the positive and the negative documents (from the context pair) as equally similar to the candidate document, there’s no direction-establishing signal.
**Feedback’s confidence is zero.**
If, from the feedback model’s perspective, both documents in a context pair are identical to the query, there’s, once again, no information on direction of more relevance.
**Context pair's confidence is zero.**
If, from the feedback model’s perspective, both documents in a context pair are identical to the query, there’s, once again, no information on direction of relevance.
In both cases the final formula should rely on the retriever's judgment, as opposed to a situation when feedback signals are strong.
One of the options to express this behaviour is through a weighted sum of retriever and feedback’s model signals.
In both cases the final formula should rely on the retriever's judgment, as opposed to a situation where feedback signals are strong.
### Naive Formula
Applying the math of edge cases and the idea that confidence and delta signals should play their individual role in the scoring, we came up with this three-parameter formula.
One of the options to express this behavior is through a weighted sum of signals, ensuring confidence and delta don't collapse into a single joint term by exponentiating one of them.
Applying the math of edge cases, we came up with this three-parameter (**a**, **b** and **c**) formula.
$$
F = a \cdot \text{score} + \sum_{i=1}^{\text{# pairs}} \text{confidence}_{i}^{b} \cdot c \cdot \text{delta}_i
F = a \cdot \text{score} + \sum_{p=1}^{\text{# pairs}} \text{confidence}_{p}^{b} \cdot c \cdot \text{delta}_p
$$
It computes the score between the query and a document-candidate on the retrieval iteration after feedback.
It computes the score between the query and a candidate document on the retrieval-with-relevance-feedback iteration.
> Is this formula the only option? Absolutely not! We’ve tried three others in experiments, and this one was the simplest working one. In future releases we even plan to allow providing custom formulas!
*As you see, the amount of context pairs used in the scoring formula can be more than one but one is also an option*.
`score`
> Is this formula set in stone? Absolutely not! We tried three others in experiments, and this one was the simplest that worked. In future releases we plan to allow providing custom formulas.
#### What Goes Into It
$\text{score}_\text{retriever}(\text{query}, \text{candidate document})$
|||
|---|---|
| What does it mean? | A similarity score between the query and the candidate embedding generated by the retriever model. For example, cosine similarity. |
| What does it mean? | A similarity score between the query and the candidate embedding generated by the retriever model. F.e., cosine similarity. |
| When is it calculated? | On the second step of retrieval, during search for more relevant candidates in the vector space. |
| Example | If $\text{cosine}(\text{query}, \text{candidate document}) = 0.83$, then $\text{score} = 0.83$. |
`confidence`
$\text{confidence}_\text{feedback}(\text{context pair}_i)$
$\text{confidence}_\text{feedback}(\text{context pair}_p)$
|||
|---|---|
| What does it mean? | A difference in relevance to the query for the two documents forming the $\text{context pair}_i$, as scored by the feedback model. |
| What does it mean? | A difference in relevance to the query for the two documents forming a context pair, as scored by the feedback model. |
| When is it calculated? | At the moment feedback is collected, right after the initial retrieval. |
| Example | We have $\text{context pair}_i$ out of $\text{doc}_1$ and $\text{doc}_2$.<br/>A feedback model scores them $0.99$ (more relevant, so “positive”) and $0.70$ (less relevant, so “negative”) respectively against the query.<br/>$\text{confidence}$ of the $\text{context pair}_i = 0.99 - 0.70 = 0.29$. |
| Example | We have $\text{context pair}_p$ out of $\text{doc}_1$ and $\text{doc}_2$.<br/>A feedback model scores them $0.99$ (more relevant, so “positive”) and $0.70$ (less relevant, so “negative”) respectively against the query.<br/>$\text{confidence}$ of the $\text{context pair}_p = 0.99 - 0.70 = 0.29$. |
`delta`
$\text{delta}_\text{retriever}(\text{context pair}_i, \text{candidate document})$
$\text{delta}_\text{retriever}(\text{context pair}_p, \text{candidate document})$
|||
|---|---|
| What does it mean? | The difference between the candidate’s similarity scores (for example, cosine similarity) to the positive document and to the negative document from the context pair. All embeddings, on which the similarity scores are computed, are generated by the retriever. |
| What does it mean? | The difference between the candidate’s similarity (e.g., cosine) to the positive document and to the negative document from the context pair. <br> All the similarity scores are calculated on embeddings generated by the retriever. |
| When is it calculated? | On the second step of retrieval, during search for more relevant candidates in the vector space. |
| Example | Given the context pair ($\text{doc}_p$, $\text{doc}_n$) and a candidate, all embedded by the retriever:<br/>If $\text{cosine}(\text{doc}_p, \text{candidate})$ and $\text{cosine}(\text{doc}_n, \text{candidate})$ are respectively $0.78$ and $0.40$.<br/>$\text{delta} = 0.78 - 0.40 = 0.38$. |
| Example | Given the context pair ($\text{doc}_1$, $\text{doc}_2$) mentioned above:<br/>If $\text{cosine}(\text{doc}_1, \text{candidate})$ and $\text{cosine}(\text{doc}_2, \text{candidate})$ are respectively $0.78$ and $0.40$.<br/>$\text{delta} = 0.78 - 0.40 = 0.38$. |
## Training Formula Parameters
Now the question is: **Where do we take a, b, and c from?**
With the right parameters a, b and c, the naive formula should rank documents better than our simple retriever.
That leads us to **minimizing Ranking Loss** as the training objective.
> We'll need to obtain formula parameters (a, b and c) through some simple training process, because score, confidence, and delta value distributions will likely vary by dataset, retriever, and feedback model.
Since the formula has only three parameters, **the amount of training data shouldn’t be big at all**. A few hundred domain relevant queries are enough.
### Training Objective
**For each query, the goal is for the feedback based scoring formula to rank documents as well as the feedback model would.**
With the right a, b and c, adjusted to dataset, retriever and feedback scores distribution, the naive formula should rank documents better than our simple retriever.
That leads us to **minimizing Pairwise Ranking Loss** as the training objective.
1. During training we initially retrieve the top X documents per query. X is chosen as large as affordable, given the cost of obtaining a golden ranking on these X documents from the feedback model.
2. The feedback model provides feedback on the top Y results, with Y (context limit) much smaller than X.
3. The remaining X - Y documents form the training pool for reranking.
Since the formula has only three parameters, **the amount of training data needed is very small**. A few hundred domain-relevant queries (per dataset) are more than enough.
For each query, **the goal is for the feedback-based scoring formula to rank documents as well as the feedback model would**.
Hence, to form a training dataset:
1. We retrieve the top X documents per query.
X is chosen as large as affordable, given the cost of obtaining a golden ranking on these X documents from the feedback model.
2. We use feedback from the top K results to mine context pairs for the formula, with K (context limit) being much smaller than X.
3. The remaining X − K documents form the training pool, on which our naive formula learns to rank better (similarly to how the feedback model would).
## Experiments
Before building a new interface, we needed to confirm our instinct: that a relevance feedback scoring can surface relevant documents that initial retrieval misses.
<TBD decide if we are including here our experiments repo>
### Setup
### Design
Each triplet **retriever**, **feedback model**, **dataset** forms one experiment.
Each triplet (retriever, feedback model, dataset) forms one experiment.
<details>
<summary><b>Datasets</b></summary>
<p>A subset of <a href="https://github.com/beir-cellar/beir">BEIR (Informational Retrieval benchmark)</a>:</p>
<ul>
<li>MSMARCO (8.84mln documents)</li>
<li>SCIDOCS (25K documents)</li>
<li>Quora (523K documents)</li>
<li>FiQA-2018 (57K documents)</li>
<li>NFCorpus (3.6K documents)</li>
</ul>
<p>— documents and queries, no qrels needed.</p>
</details>
**Dataset**
A subset of [BEIR (Informational Retrieval benchmark)](https://github.com/beir-cellar/beir): MSMARCO (8.84mln documents), SCIDOCS (25K documents), Quora (523K documents), FiQA-2018 (57K documents) and NFCorpus (3.6K documents) – documents and queries, no qrels needed.
<details>
<summary><b>Retrievers</b></summary>
<p>As retrievers we chose:</p>
<ul>
<li><a href="https://huggingface.co/jinaai/jina-embeddings-v2-base-en">jina-embeddings-v2-base-en</a> (768 output dimensions)</li>
<li><a href="https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1">mxbai-embed-large-v1</a> (1024 output dimensions)</li>
<li><a href="https://huggingface.co/michaelfeil/Qwen3-Embedding-0.6B-auto">Qwen3-Embedding-0.6B</a> (1024 output dimensions)</li>
</ul>
</details>
**Retriever**
As retrievers we chose [jina-embeddings-v2-base-en](https://huggingface.co/jinaai/jina-embeddings-v2-base-en) (768 output dimensions), [mxbai-embed-large-v1](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) (1024 output dimensions) and [Qwen3-Embedding-0.6B](https://huggingface.co/michaelfeil/Qwen3-Embedding-0.6B-auto) (1024 output dimensions).
<details>
<summary><b>Feedback models</b></summary>
<p>The cheapest feedback type suitable for testing the hypothesis that came to mind was <b>using embedding models (bi-encoders) with higher dimensionality than those used for retrieval</b>.</p>
<p>The feedback from these models is a similarity score between query and document embeddings (so, simply put, cosine similarity).</p>
<p>We chose as feedback models:</p>
<ul>
<li><a href="https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1">mxbai-embed-large-v1</a> (1024 output dimensions)</li>
<li><a href="https://huggingface.co/michaelfeil/Qwen3-Embedding-0.6B-auto">Qwen3-Embedding-0.6B</a> (1024 output dimensions)</li>
<li><a href="https://huggingface.co/michaelfeil/Qwen3-Embedding-4B-auto">Qwen3-Embedding-4B</a> (2560 output dimensions)</li>
<li><a href="https://huggingface.co/colbert-ir/colbertv2.0">colBERTv2.0</a> (multivectors of 128 dimensions each)</li>
</ul>
</details>
**Feedback model**
The cheapest feedback model that came to mind was using **embedding models with higher dimensionality than the ones used for retrieval**.
> A feedback model can be anything: a bigger embedding model, a cross encoder, a custom ranking model, an LLM, or an agent.
The pointwise feedback from these models is simply a similarity score between query and document embeddings (so, in most cases, cosine similarity).
We chose as feedback models [mxbai-embed-large-v1](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) (1024 output dimensions), [Qwen3-Embedding-0.6B](https://huggingface.co/michaelfeil/Qwen3-Embedding-0.6B-auto) (1024 output dimensions), [Qwen3-Embedding-4B](https://huggingface.co/michaelfeil/Qwen3-Embedding-4B-auto) (2560 output dimensions), and [colBERTv2.0](https://huggingface.co/colbert-ir/colbertv2.0) (multivectors of 128 dimensions each).
With Qdrant and [Qdrant’s Cloud Inference](https://qdrant.tech/cloud-inference/) at our disposal, we ran embedding inference once for all datasets on both retrieval and feedback models, stored the embeddings in Qdrant, and then experimented with different formulas. This way, we could even use in our experiments one model interchangeably as a retriever and a feedback model.
With [Qdrant’s Cloud Inference](https://qdrant.tech/cloud-inference/) at our disposal, we ran embedding inference once for all datasets on both retrieval and feedback models, stored the embeddings in Qdrant, and then experimented with different formulas.
This way, we could use in our experiments one model interchangeably as a retriever and a feedback model.
### Metric
We wanted to show that relevance feedback-based scoring increases recall of the retrieved results. That means **surfacing documents more relevant to the query than initial retrieval results**.
We want to check if relevance feedback-based retrieval increases relevance recall. That means surfacing documents more relevant to the query than the vanilla retriever did.
**What we’re measuring**
Does our feedback-based scoring formula pull more relevant documents up into the evaluation window N than the baseline retriever?
**How did we emulate "we-want-to-surface" quality in experiments?**
**Set up:**
* For each query, the retriever returned a ranked list of documents.
* We obtained a ground-truth relevance scoring of this list using the feedback model.
* We mined context pairs for the feedback-based scoring formula only from the top-K retrieved results.
This top-K emulated the context limit in production, aka the initial retrieval results available for feedback.
* From the feedback model's scores within the top-K, we took the highest score as a **threshold**.
* For each query, the retriever returns a ranked list from the initial retrieval.
* The feedback model sees only the context limit of results (say, top-X, X being 3-5). We mine our context pairs and their confidence.
* We set a **threshold**. It is the best relevance score between query and document within that context limit as assigned by the feedback model.
Any document in the list **outside the top-K** whose feedback model score **exceeded this threshold** was considered a **desired result** for "the next retrieval iteration".
Any document **outside the context limit** that **would score above the threshold** is a “we-want-to-surface-it” item.
{{< figure src="/articles_data/relevance-feedback/goal.png" alt="Diagram contrasting retriever's ranking with ranking made by a feedback model. The top section shows the retriever model’s initial ranking, where only the highest scoring items on the left are included in the default top K. The bottom section shows the golden ground-truth relevance scoring on the same cadidates from a feedback model, with some different items now receiving higher scores. Orange bars labeled Desired results appear below the original top K threshold line, illustrating documents that were scored too low by the retriever. The figure demonstrates more relevant documents that were not included in the original top K results that we'd want to surface." caption="More relevant documents we'd like to surface">}}
{{< figure src="/articles_data/relevance-feedback/metric_goal.png" alt="The image shows a vertical list of retrieved documents represented as colored blocks, coloured based on their relevance in shades of red, green and gray. They're under the heading ‘retriever.’ A dotted line labeled ‘context limit’ indicates which top items are used for feedback. Below this line are additional retrieved documents, including blocks in deeper green shades. Arrows point from these deeper-green blocks to text reading ‘we want to surface,’ indicating that these items are more strongly positive (greener) than the greenest documents within the context limit and therefore should be surfaced higher in the ranking." caption="\"We-want-to-surface-it\" items -- outside the context limit and more relevant than any documents within it" width="80%" >}}
**What we wanted to measure**
On the next retrieval iteration, can our feedback-based formula pull more relevant documents into the top N than the vanilla retriever did?
*Relevance is judged by the feedback model's ground-truth scores.*
**Compare baseline retriever versus feedback-based scoring with the custom metric@N:**
For that, we came up with the **abovethreshold@N** metric.
* Look **at the next N positions (evaluation window)** after the top-X documents used for feedback (we're excluding context from evaluation).
* **Baseline:** count how many “we-want-to-surface-it” documents already appear there in the baseline ranking.
* **Rescored:** rescore and rerank documents, excluding ones in the context, using a trained formula. Check the first N positions (evaluation window) of the new ranking and count again the “we-want-to-surface-it” items.
#### Abovethreshold@N
{{< figure src="/articles_data/relevance-feedback/metric_at_N.png" alt="The image compares a baseline retriever with feedback-based scoring within a shared evaluation window. Inside a dashed rectangle labeled ‘evaluation window,’ two columns of colored blocks are shown: the left column for the retriever and the right column for feedback-based scoring. The blocks vary in color intensity, with deeper green shades indicating higher relevance. The diagram highlights how many highly relevant ‘we-want-to-surface-it’ items (shown in deeper green) appear inside the evaluation window before and after applying feedback-based scoring and reranking." caption="Measuring how many \"we-want-to-surface-it\" items are within the evaluation window of size N">}}
A custom metric to compare vanilla retriever versus relevance feedback-based scoring.
Sum counts across all queries for baseline and for rescored and calculate the **relative gain**.
**How we computed it (per query):**
We looked at N positions after the top-K documents used for mining context pairs (positions K+1 through K+N).
* **Vanilla retriever:** counted how many "we-want-to-surface" documents already appeared in these positions.
* **Relevance feedback-based scoring:** rescored all remaining documents (outside top-K) using a trained naive formula, reranked them, and counted how many "we-want-to-surface" documents appeared in the first N positions of the new ranking.
### Experiment Parameters
{{< figure src="/articles_data/relevance-feedback/metric.png" alt="Diagram showing the abovethreshold at 10 metric comparing vanilla second retrieval with feedback-based second retrieval. The threshold is the highest feedback model score within the original top K. Documents outside the top K whose feedback score exceeds this threshold are considered desired results. In the vanilla ranking, zero out of two desired documents appear in the next 10 positions, while after feedback-based rescoring, two out of two appear, showing improved recall of relevant documents." caption="Abovethreshold@N metric">}}
To get one number per test set, we summed abovethreshold@N counts across all queries per method and calculated the **relative gain**: `(feedback_count - vanilla_count) / vanilla_count`.
### Parameters
<details>
<summary><b>Training data size</b></summary>
@@ -321,8 +353,8 @@ Queries for training are split in 50% train, 50% validation.
| Parameter | Value |
|-------------------------------------------------------------|------:|
| Golden Ranking Limit (documents per query) | 100 |
| Context Limit (to mine context pairs with the feedback model)| 5 |
| Golden feedback model scoring limit (documents per query) | 100 |
| Context Limit, aka top K (to mine context pairs) | 5 |
| Context pairs used (out of all, sorted by the confidence) | top-1 |
| Learning rate | 0.005 |
| Epochs with early stopping and patience 200 | 2000 |
@@ -334,19 +366,19 @@ Queries for training are split in 50% train, 50% validation.
| Parameter | Value |
|-------------------------------------------------------------|------:|
| Context Limit (to mine context pairs with the feedback model)| 3 |
| Context Limit, aka top K (to mine context pairs) | 3 |
| Context pairs used (out of all, sorted by the confidence) | all |
| Evaluation window size (metric@N) | 10 |
| Evaluation window size (N in abovethreshold@N) | 10 |
</details>
### Results
Rescoring on the client side humongous datasets, such as MSMARCO, for every query would not have been fun, especially while trying different formulas and hyperparameters. We faced the same problem that limits many relevance feedback researchers and users, making them test approaches only on a subset of all documents.
Rescoring humongous datasets like MSMARCO on the user side for every query would not have been fun, especially while iterating on different formulas and hyperparameters. We faced the same problem that limits many relevance feedback researchers and practitioners, testing approaches only on a subset of all documents.
Our relevance feedback interface is a remedy against this limitation, but first, we needed experiment results to further justify the interface's implementation. So, this project started turning into a chicken-egg problem. **Hen**ce, to break the loop, in the initial experiments we simulated feedback-based scoring on a limited subset of documents per query -- **100**. We decided to use only **three** documents as a context limit.
Our Relevance Feedback Query API is a remedy against this limitation, but first we needed experiment results to justify its implementation. A chicken-and-egg problem. **Hen**ce, to break the loop, we simulated feedback-based scoring on a hundred documents per query.
Out of all pairs of retriever and feedback models, we got three leaders with the following relative gain in **metric@10** compared to the baseline retriever.
Out of all retriever–feedback model pairs, three leaders emerged with the following relative gain in **abovethreshold@10** compared to the vanilla retriever:
| | Qwen3-0.6B → colBERTv2.0 | Qwen3-0.6B → Qwen3-4B | mxbai-large-v1 → colBERTv2.0 |
| ----- | ----- | ----- | ----- |
@@ -356,7 +388,7 @@ Out of all pairs of retriever and feedback models, we got three leaders with the
| **MSMARCO** | **+23.23%** | +16.73% | +2.40% |
| **Quora** | **+5.04%** | +2.67% | 0.00% |
[jina-embeddings-v2-base-en](https://huggingface.co/jinaai/jina-embeddings-v2-base-en), as a smaller, and, consequently, less expressive retriever, was not responsive to the feedback. Results are noisy, with no noticeable improvement from adding the feedback signal.
[jina-embeddings-v2-base-en](https://huggingface.co/jinaai/jina-embeddings-v2-base-en), being a smaller and less expressive retriever, did not benefit from the feedback signal. Results varied significantly across queries, with no consistent improvement from adding feedback-based scoring.
| | jina-v2-base → mxbai-large-v1 | jina-v2-base → Qwen3-0.6B | jina-v2-base → Qwen3-4B |
| ----- | ----- | ----- | ----- |
@@ -366,9 +398,33 @@ Out of all pairs of retriever and feedback models, we got three leaders with the
| **MSMARCO** | +2.57% | +2.23% | **+2.82%** |
| **Quora** | 0.00% | −1.37% | 0.00% |
#### Using Entire Vector Space
#### Takeaways
The results above justified our relevance-feedback approach and FeedbackQuery implementation. After completing the latter, we checked how our naive formula rescores the **entire** SCIDOCS dataset using the identical testing parameters.
* The retriever's expressiveness limits how much it can benefit from feedback. The retriever operates in a lower-dimensional space and can't capture all the distinctions the feedback model makes. Past a certain point, a more sophisticated feedback model won't help.
* Relevance feedback scoring is most effective when the feedback model (reasonably) disagrees with the retriever's ordering within the top-K (context limit). "Reasonably" meaning the feedback aligns with the user's actual notion of relevance.
* Initially we used only the single highest-confidence context pair in the feedback-based scoring formula, both during training and at inference. We then discovered that at inference time, incorporating signals from additional (all) context pairs improves results.
## Relevance Feedback Query
The results above were convincing enough for us to justify Relevance Feedback Query implementation.
### How to Use Relevance Feedback Query
The method, like everything in Qdrant, is open source. We did a python package that computes for you customized Naive Formula weights for your dataset, retriever, and feedback model <TBD link to PyPi>.
What you need is a Qdrant collection, an idea of which feedback model would you like to use with your retriever and, optionally, a set of collection-specific queries.
> A feedback model can be anything: a bi-encoder, a cross-encoder or an LLM.
After you've obtained the weights, simply plug them in and [run Relevance Feedback Query with your favourite Qdrant client](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback).
### Evaluation
Additionally, framework provides you with an evaluation module, implementing two metrics: **relative gain** based on the **abovethreshold@N** metric from experiments and, additionally, we've a metric which is less custom & more recognizeable for people in search & more related to ranking **Discounted cumulative gain (DCG) Win Rate**.
We checked how our naive formula rescores the **entire** SCIDOCS dataset using the identical testing parameters.
| | Qwen3-0.6B → colBERTv2.0 | Qwen3-0.6B → Qwen3-4B | mxbai-large-v1 → colBERTv2.0 |
| ----- | ----- | ----- | ----- |
@@ -383,15 +439,17 @@ The results above justified our relevance-feedback approach and FeedbackQuery im
</details>
#### Takeaways:
**Discounted cumulative gain (DCG) Win Rate**
* The retriever’s expressiveness limits how much it can respond to feedback. Moreover, past a certain point, making the feedback model more sophisticated will not help, because the retriever works in a lower dimensional space and often can’t capture the distinctions the feedback model is making.
* Initially we used only one context pair with the highest confidence in the feedback-based scoring formula, both when training and when applying feedback based scoring. Then we discovered that for the latter, adding signals from other pairs improves results. They are added as a summand per pair: $+ \text{confidence}(\text{context\_pair}_i)^b \cdot c \cdot \text{delta}(\text{document\_candidate}, \text{context\_pair}_i)$
This metric measures the percentage of queries where one retrieval method outperforms another in terms of the Discounted Cumulative Gain (DCG) of the retrieved responses.
The relevance of responses is determined by the feedback model scores.
For each query, we compute DCG@N for both methods. The method with the higher DCG gets a "win".
If DCG values are equal, it is counted as a "tie".
## Conclusion
We’ve released a new relevance feedback tool in Qdrant 1.17.0 <TBD link> to help increasing the relevance of vector search results. **It is built for scale, it’s cheap, customizable and universal.**
We’ve released [a new relevance feedback tool in Qdrant 1.17.0](https://qdrant.tech/blog/qdrant-1.17.x/#relevance-feedback-query) to help increasing the relevance of vector search results. **It is built for scale, it’s cheap, customizable and universal.**
**It is cheap to use**, because the time and resources spent on getting relevance feedback are minimal.
@@ -403,13 +461,12 @@ It uses **the whole vector space of documents** to search for more relevant resu
It’s a tool for production.
### How to Use Relevance Feedback
### When to Use Relevance Feedback
​​Use the relevance feedback method once basic retrieval is in place and you are looking for extra instruments to boost result relevance, such as hybrid search, query rewriting, reranking, or score boosting. (If you are not looking for them, why not? There is no limit to perfection!)
**It’s here not to replace but to complement other search relevance tools.** For example, it is a good aid for your search agents, letting you propagate their use case understanding directly to the vector search index.
The method, like everything in Qdrant, is open source. We have provided a framework that lets you compute customized FeedbackQuery weights for your dataset, retriever, and feedback model<TBD link>.
If you would like additional advice on FeedbackQuery setup in production or have ideas to enhance the method, please write to us<TBD some discord channel>.
If you would like additional advice on FeedbackQuery setup in production or have ideas to enhance the method, talk to us [Office Hours, for example].