Files
landing_page/qdrant-landing/content/documentation/concepts/hybrid-queries.md
T
378837432e Publish docs for version 1.17 (#2137)
* Cluster Telemetry docs (#2068)

* mention cluster telemetry

* small reword

* Docs for read_fan_out_delay_ms

* Add warning against setting threshold too low

* Small docs change for inference API keys

* Trigger Build

* Add guide with tips for low latency search (#2136)

* Update frontmatter weights

* Add 'Tips for Low-Latency Search' guide

* Review feedback

* Temporarily bump Rust client to v1-17-upgrade branch

* Docs for audit logging (#2141)

* Docs for audit logging

* Consistent title casing

* Review feedback

* Docs for optimization monitoring (#2121)

* Docs for optimization monitoring

* Review feedback

* Optimization monitoring is cluster-wide now

* Upgrade code snippet checker to 1.17

* Fix broken Python snippets

* Relevance Feedback docs (#2060)

* add relevance feedback in Explore page

* Review

* Trigger Build

* Trigger Build

* Create new 'Search Relevance' concept page

* Tweaks

* Update links to moved content

* add rust snippet

* TS anippets

* Clarification about using point IDs

* Add links

* Restructure paragraphs

* docs: Go snippet

Signed-off-by: Anush008 <mail@anush.sh>

* docs: Missed Java snippets with C#

Signed-off-by: Anush008 <mail@anush.sh>

* new: add python snippets

* Make Java and Rust snippets testable

* Remove unnecessary styling

---------

Signed-off-by: Anush008 <mail@anush.sh>
Co-authored-by: Evgeniya Sukhodolskaya <suxodolskaya97@gmail.com>
Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
Co-authored-by: Anush008 <mail@anush.sh>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>

* Docs for enable_hnsw (#2080)

* Move 'Filterable HNSW Index' section under 'Vector Index'

* Add docs for enable_hnsw

* docs: Go snippets

Signed-off-by: Anush008 <mail@anush.sh>

* docs: Java snippets

Signed-off-by: Anush008 <mail@anush.sh>

* docs: C# snippets

Signed-off-by: Anush008 <mail@anush.sh>

* Add Rust snippet

* TS snippets

* new: add python snippets

---------

Signed-off-by: Anush008 <mail@anush.sh>
Co-authored-by: Anush008 <mail@anush.sh>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>

* Docs for list shard keys API (#2082)

* Docs for list shard keys

* docs: Go snippets

Signed-off-by: Anush008 <mail@anush.sh>

* docs: Java snippets

Signed-off-by: Anush008 <mail@anush.sh>

* docs: C# snippets

Signed-off-by: Anush008 <mail@anush.sh>

* Add Rust snippet

* docs: TS snippets

* new: add python snippets

---------

Signed-off-by: Anush008 <mail@anush.sh>
Co-authored-by: Anush008 <mail@anush.sh>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>

* Docs for update_mode (#2097)

* Docs for update_mode

* docs: C#, Go, Java snippets

Signed-off-by: Anush008 <mail@anush.sh>

* doc: Remove _ from C# snippet

Signed-off-by: Anush008 <mail@anush.sh>

* Add Rust snippet

* Fix some snippets

* Review feedback

* ts snippets

* new: add python snippets

* fix: add generated python.md

---------

Signed-off-by: Anush008 <mail@anush.sh>
Co-authored-by: Anush008 <mail@anush.sh>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>

* Docs for weighted RRF (#2132)

* Docs for weighted RRF

* docs: Go, Java, C# snippets

Signed-off-by: Anush008 <mail@anush.sh>

* Add Rust snippet

* Apply suggestions from code review

Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>

* Delete landing_page.sln

* Review feedback

* ts snippets

* new: add python snippets

* Trigger Build

---------

Signed-off-by: Anush008 <mail@anush.sh>
Co-authored-by: Anush008 <mail@anush.sh>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>

* Clean up Go and Java snippets

* Revert go snippet change; target Go client 1.17.1

* Updated Go snippet

* Trigger Build

* Update Rust lockfile

---------

Signed-off-by: Anush008 <mail@anush.sh>
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
Co-authored-by: Daniel Boros <56868953+dancixx@users.noreply.github.com>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Evgeniya Sukhodolskaya <suxodolskaya97@gmail.com>
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
Co-authored-by: Anush008 <mail@anush.sh>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
2026-02-20 10:42:12 +01:00

7.5 KiB

title, weight, aliases, hideInSidebar
title weight aliases hideInSidebar
Hybrid Queries 57
../hybrid-queries
false

Hybrid and Multi-Stage Queries

Available as of v1.10.0

With the introduction of multiple named vectors per point, there are use-cases when the best search is obtained by combining multiple queries, or by performing the search in more than one stage.

Qdrant has a flexible and universal interface to make this possible, called Query API (API reference).

The main component for making the combinations of queries possible is the prefetch parameter, which enables making sub-requests.

Specifically, whenever a query has at least one prefetch, Qdrant will:

  1. Perform the prefetch query (or queries),
  2. Apply the main query over the results of its prefetch(es).

Additionally, prefetches can have prefetches themselves, so you can have nested prefetches.

One of the most common problems when you have different representations of the same data is to combine the queried points for each representation into a single result.

{{< figure src="/docs/fusion-idea.png" caption="Fusing results from multiple queries" width="80%" >}}

For example, in text search, it is often useful to combine dense and sparse vectors to get the best of both worlds: semantic understanding from dense vectors and precise word matching from sparse vectors.

Qdrant has a few ways of fusing the results from different queries: rrf and dbsf

Reciprocal Rank Fusion (RRF)

RRF considers the positions of results within each query and boosts those that appear closer to the top in multiple sets of results. The score of a document is calculated using its rank in each result set: score(d\in D) = \sum_{r_d\in R(d)} \frac{1}{k + \frac{r_d + 1}{w_r} - 1}

Where:

  • D the set of points across all results
  • R(d) is the set of rankings for a particular document
  • k is a constant (set to 2 by default)
  • r is an ordered set of results from one source
  • r_d is the rank of document d in ranking r
  • w_r is the weight of ranking r (set to 1 by default)

Because w_r defaults to 1, without setting explicit weights, the formula can be simplified to the original RRF function:

score(d\in D) = \sum_{r_d\in R(d)} \frac{1}{k + r_d}

Here is an example of RRF for a query containing two prefetches against different named vectors configured to hold sparse and dense vectors, respectively.

{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rrf/" >}}

Setting RRF Constant k

Available as of v1.16.0

To change the value of constant k in the formula, use the dedicated rrf query.

{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rrf-k/" >}}

Weighted RRF

Available as of v1.17.0

By default, each query is assigned an equal weight. In reality, some queries are stronger, more discriminative, or more domain-specific than others. For example, a semantic search model understands meaning better than a simple keyword matcher. Assigning equal weight to both can cause the weaker model to negatively influence results, leading to a suboptimal search experience. To address this, you can assign greater weight to rankers that perform well.

The rrf query allows you to configure relative weights for each of the prefetches. For example, if you have two prefetches and assign a weight of 3.0 to the first and 1.0 to the second, a document ranked third in the first query scores the same as a document ranked first in the second query. In the case of non-overlapping result sets, these weights return three results from the first set for every one result from the second set.

Weights should be provided as an array of numbers, where each weight is applied to the corresponding prefetch in the order they are defined. The number of weights must match the number of prefetches.

{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rrf-weights/" >}}

Distribution-Based Score Fusion (DBSF)

Available as of v1.11.0

DBSF normalizes the scores of the points in each query, using the mean +/- the 3rd standard deviation as limits, and then sums the scores of the same point across different queries.

Multi-stage queries

In general, larger vector representations give more accurate search results, but makes them more expensive to compute.

Splitting the search into two stages is a known technique to mitigate this effect:

  • First, use a smaller and cheaper representation to get a large list of candidates.
  • Then, re-score the candidates using the larger and more accurate representation.

There are a few ways to build search architectures around this idea:

  • The quantized vectors as a first stage, and the full-precision vectors as a second stage.
  • Leverage Matryoshka Representation Learning (MRL) to generate candidate vectors with a shorter vector, and then refine them with a longer one.
  • Use regular dense vectors to pre-fetch the candidates, and then re-score them with a multi-vector model like ColBERT.

To get the best of all worlds, Qdrant has a convenient interface to perform the queries in stages, such that the coarse results are fetched first, and then they are refined later with larger vectors.

Re-scoring examples

Fetch 1000 results using a shorter MRL byte vector, then re-score them using the full vector and get the top 10.

{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring/" >}}

Fetch 100 results using the default vector, then re-score them using a multi-vector to get the top 10.

{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring-multivector/" >}}

It is possible to combine all the above techniques in a single query:

{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring-multistage/" >}}

Grouping

Available as of v1.11.0

It is possible to group results by a certain field. This is useful when you have multiple points for the same item, and you want to avoid redundancy of the same item in the results.

REST API (Schema):

{{< code-snippet path="/documentation/headless/snippets/query-groups/basic/" >}}

For more information on the grouping capabilities refer to the reference documentation for search with grouping and lookup.