Feedback models
The cheapest feedback type suitable for testing the hypothesis that came to mind was using embedding models (bi-encoders) with higher dimensionality than those used for retrieval.
The feedback from these models is a similarity score between query and document embeddings (so, simply put, cosine similarity).
We chose as feedback models:
With [Qdrant’s Cloud Inference](https://qdrant.tech/cloud-inference/) at our disposal, we ran embedding inference once for all datasets on both retrieval and feedback models, stored the embeddings in Qdrant, and then experimented with different formulas.
This way, we could use in our experiments one model interchangeably as a retriever and a feedback model.
### Metric
We want to check if relevance feedback-based retrieval increases relevance recall. That means surfacing documents more relevant to the query than the vanilla retriever did.
**How did we emulate "we-want-to-surface" quality in experiments?**
* For each query, the retriever returned a ranked list of documents.
* We obtained a ground-truth relevance scoring of this list using the feedback model.
* We mined context pairs for the feedback-based scoring formula only from the top-K retrieved results.
This top-K emulated the context limit in production, aka the initial retrieval results available for feedback.
* From the feedback model's scores within the top-K, we took the highest score as a **threshold**.
Any document in the list **outside the top-K** whose feedback model score **exceeded this threshold** was considered a **desired result** for "the next retrieval iteration".
{{< figure src="/articles_data/relevance-feedback/goal.png" alt="Diagram contrasting retriever's ranking with ranking made by a feedback model. The top section shows the retriever model’s initial ranking, where only the highest scoring items on the left are included in the default top K. The bottom section shows the golden ground-truth relevance scoring on the same cadidates from a feedback model, with some different items now receiving higher scores. Orange bars labeled Desired results appear below the original top K threshold line, illustrating documents that were scored too low by the retriever. The figure demonstrates more relevant documents that were not included in the original top K results that we'd want to surface." caption="More relevant documents we'd like to surface">}}
**What we wanted to measure**
On the next retrieval iteration, can our feedback-based formula pull more relevant documents into the top N than the vanilla retriever did?
*Relevance is judged by the feedback model's ground-truth scores.*
For that, we came up with the **abovethreshold@K** metric.
#### Abovethreshold@N
A custom metric to compare vanilla retriever versus relevance feedback-based scoring.
**How we computed it (per query):**
We looked at N positions after the top-K documents used for mining context pairs (positions K+1 through K+N).
* **Vanilla retriever:** counted how many "we-want-to-surface" documents already appeared in these positions.
* **Relevance feedback-based scoring:** rescored all remaining documents (outside top-K) using a trained naive formula, reranked them, and counted how many "we-want-to-surface" documents appeared in the first N positions of the new ranking.
{{< figure src="/articles_data/relevance-feedback/metric.png" alt="Diagram showing the abovethreshold at 10 metric comparing vanilla second retrieval with feedback-based second retrieval. The threshold is the highest feedback model score within the original top K. Documents outside the top K whose feedback score exceeds this threshold are considered desired results. In the vanilla ranking, zero out of two desired documents appear in the next 10 positions, while after feedback-based rescoring, two out of two appear, showing improved recall of relevant documents." caption="Abovethreshold@N metric">}}
To get one number per test set, we summed abovethreshold@N counts across all queries per method and calculated the **relative gain**: `(feedback_count - vanilla_count) / vanilla_count`.
### Parameters